EDBT 2026 Demo / reviewers in the wild / expert
Chao Yang 0021
dblp:00/5867-21
· DBLP profile ↗
43ranked-venue papers
6as first author
27since 2021 · last 2026
0000-0001-9276-5673ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 37 · 5 first-author · 22 since 2021Artificial intelligence and machine learning · 5 · 1 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | VIQ-360: A New Viewport-Based Omnidirectional Image Quality Assessment Database
Xiangxu Yu, Chao Yang 0021, Xinpeng Huang, Ping An 0001 |
QoMEX | 3 |
| 2026 | A Subjective Quality Database for Human-AI Co-Created Images
Xiangxu Yu, Jiyan Tong, Chao Yang 0021, Xinpeng Huang, Ping An 0001 |
QoMEX | 5 |
| 2026 | DMGNet: Discriminative multi-view geometry learning with hybrid-domain enhancement for light field occlusion removal
Jieyu Chen, Ping An 0001, Xinpeng Huang, Chao Yang 0021, Ce Zhu, Sanghoon Lee 0001 |
Expert Syst. Appl. | 4 |
| 2026 | Enhancing joint human-machine image compression via Chebyshev space modulation
Zhicheng Ma, Ping An 0001, Shipei Wang, Chao Yang 0021, Xinpeng Huang |
Multim. Syst. | 4 |
| 2026 | Generic feature extraction and compression for human and machine-oriented vision
Kunqiang Huang, Ping An 0001, Chao Yang 0021, Shipei Wang, Xinpeng Huang, Liquan Shen |
Signal Process. | 3 |
| 2026 | Viewport-Patch Extraction Enhanced 360$^\circ$ Video Quality AssessmentabstractWith the rising adoption of 360$^\circ$video in virtual reality (VR) applications, assessing its perceptual quality remains a challenge due to projection-induced distortions in equirectangular projection (ERP) formats. Traditional sliding-window cropping methods often distort high-latitude content and fail to reflect the actual viewing experience. To address this, we propose a novel viewport patch-based video quality assessment (VQA) method. By sampling view directions on the sphere and applying gnomonic projection, our method extracts undistorted and perceptually consistent viewport patches that preserve both spatial fidelity and full-frame coverage. We further design a two-stream network that jointly models high-frequency distortion and residual information over time, enhanced by squeeze-and-excitation (SE) attention to capture spatial-temporal features. Experiments and analysis show that our method significantly improves the accuracy and reliability of 360$^\circ$VQA, achieving PLCC/SROCC values of 0.9603/0.9628 on the VQA-ODV dataset and 0.9585/0.9400 on the BIT360 dataset, with only 0.22M parameters. Code is available athttps://github.com/yeonhw/VP-VQA. Chao Yang 0021, Ping An 0001, Xinpeng Huang |
IEEE Signal Process. Lett. | 2 |
| 2026 | Low-Bitrate Light Field Video Compression Through Key Sequences Encoding and Joint Reconstruction NetworkabstractLight field (LF) videos contain rich spatial, angular, and temporal information, resulting in immense data volumes and posing significant challenges for low-bitrate compression. Existing LF video compression methods focus on modifying the structure of traditional video codecs to encode all LF views, but they are insufficient to achieve low-bitrate compression of LF video. To address these limitations, we propose a low-bitrate LF video compression framework that exploits spatial-angular-temporal correlations through sparse coding and joint reconstruction. On the encoding side, we introduce a content-adaptive prediction structure for sparse key view sequences selection. This structure is adapted to LF video content, leveraging the most similar view as a reference to enhance prediction accuracy and significantly reduce bitrate. On the decoding side, we observe that pixels missing in the current view are often captured in adjacent angular and/or temporal views. As a result, we develop a spatial-angular-temporal based joint reconstruction network that integrates cues across the different domains. This approach supplements missing texture details near occlusion areas and reconstructs high-quality non-key views. Experimental results demonstrate the efficiency of our framework, achieving an average gain of about 60 % in terms of bitrate savings and 2 dB in terms of reconstruction quality compared to the state-of-the-art methods. Xinpeng Huang, Chao Yang 0021, Mounir Kaaniche, Qiuwen Zhang, Ping An 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | Layered and scalable image coding with semantic features for human and machine
Jiao Wei, Ping An 0001, Shipei Wang, Kunqiang Huang, Chao Yang 0021, Xinpeng Huang |
Eng. Appl. Artif. Intell. | 5 |
| 2025 | Mask-Aware Light Field De-Occlusion With Gated Feature Aggregation and Texture-Semantic AttentionabstractA light field image records rich information of a scene from multiple views, thereby providing complementary information for occlusion removal. However, current occlusion removal methods have several issues: 1) inefficient exploitation of spatial and angular complementary information among views; 2) indistinguishable treatment of pixels from foreground occlusion and background; and 3) insufficient exploration of spatial detail supplementation. Therefore, in this article, we propose a mask-aware de-occlusion network (MANet). Specifically, MANet is a joint training network that integrates the occlusion mask predictor (OMP) and the occlusion remover (OR). First, OMP is proposed to provide the location of occluded regions for OR, as the occlusion removal task is ill-posed without occluded region localization. In OR, we introduce gated spatial-angular feature aggregation, which uses a soft gating mechanism to focus on spatial-angular interaction features in non-occluded regions, extracting effective aggregated features specific to the de-occlusion. Then, we design a complementary strategy to fully utilize spatial-angular information among views. Finally, we propose texture-semantic attention to improve the performance of detail generation. Experimental results demonstrate the superiority of MANet, with substantial improvements in both PSNR and SSIM metrics. Moreover, MANet stands out with an efficient parameter count of 2.4 M, making it a promising solution for real-world applications in public safety and security surveillance. Jieyu Chen, Ping An 0001, Xinpeng Huang, Chao Yang 0021, Liquan Shen |
IEEE Trans. Multim. | 5 |
| 2025 | Feature Quality Assessment: A Database and A Lightweight Objective MethodabstractIn the era of Artificial Intelligence, visual data gathered by edge devices could be primarily utilized for machine vision tasks. The prominent coding frameworks accomplish this by extracting and compressing features extracted from input data. As such, the quality of these features is vital, as they reflect the performance of the coding framework. However, much less work has been dedicated to quality assessment on features, impeding the optimization of the coding system. In this work, we pioneer to explore the feature quality assessment by creating a novel database tailored for features, with the quality ground-truth for each feature. Then, we propose a lightweight feature quality assessment method, called Lightweight Feature Quality Assessment (LFQA). We analyze the feature characteristics from the perspective of spatial and channel thoroughly, and the framework of LFQA is designed based on the analysis results. Experimental results demonstrate that LFQA accurately evaluates the quality of features, reaching a notable Spearman Rank-Order Correlation Coefficient of 85.38%, and exhibits competitive performance in improving the performance of video coding for machine system. Furthermore, LFQA has fewer model parameters and faster inference speed, ensuring a wide range of promising applications. Shipei Wang, Ping An 0001, Chao Yang 0021, Gongyang Li, Xinpeng Huang, Shiqi Wang 0001 |
IEEE Trans. Multim. | 3 |
| 2025 | Dual-Guided Video Frame Interpolation With Spatial-Temporal Global AttentionabstractVideo frame interpolation technology improves visual experience with the development of deep learning. However, capturing large motions while synthesizing fine texture details remains a challenging task. Regarding large motion scenarios, some pioneering Transformer-based methods primarily rely on local attention, which does not fully leverage the global receptive field advantage. To address this issue, this paper proposes to further broaden the receptive field of the Transformer to capture more correlations in the video frame interpolation task. Specifically, we propose a global self-attention mechanism in the form of spatial-temporal separation. Regarding texture details, since roughly enlarging the receptive field results in the loss of details, we propose to use large motion information in both feature and pixel spaces as a dual-guided prior to enhance detail synthesis. The separable attention mechanism and the straightforward frame synthesis design significantly enhance the resource efficiency of our model. Extensive experiments show that our method achieves state-of-the-art performance, effectively capturing large motions and preserving texture details. Baojun Zhou, Xinpeng Huang, Gongyang Li, Chao Yang 0021, Liquan Shen, Ping An 0001 |
IEEE Trans. Multim. | 4 |
| 2025 | Towards 360 VR Sickness Mitigation: From Virtual Reality Eye-Tracking to Visual CommunicationabstractMost 360 virtual reality (VR) contents have been developed without considering that users could be affected by VR sickness. Accordingly, users' viewing safety has been steadily highlighted as a critical problem in the VR market. In this study, we investigate a novel VR sickness mitigation framework based on human visual characteristics for the rendered VR content. First, we build a large-scale 360 VR content database termed VRSP360 (VR Sickness and Presence 360) dedicated to the analysis of VR sickness and thoroughly conduct eye-tracking experiments to measure human perception. In the experiment, we observe that the users' gaze distribution is highly center-biased when they experience excessive VR sickness. From this observation, we design a foveated filtering framework that limits high-frequency textures in the peripheral view to mitigate VR sickness. Particularly, given the human visual system's (HVS) non-uniform resolution with respect to the fovea, we also adopt the foveation-based filtering method using the trade-off between sickness mitigation and presence conservation, which reduces any loss in perceptual quality despite the filtering. We further demonstrate that our framework can effectively compress visual information by applying foveated compression. In addition, we develop two metrics (visual texture index and perceptual information index) to measure the effective preservation of user-perceived information despite the filtration of peripheral vision textures by our proposed mitigation method. Through rigorous subjective evaluation on both original content and its VR-sickness-mitigated version, we demonstrate that the proposed framework successfully mitigates VR sickness with a reduction rate of $\sim$∼19% on the proposed dataset. Jeonghaeng Lee, Woojae Kim, Chao Yang 0021, Ping An 0001, Sanghoon Lee 0001 |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2024 | Adaptive Threshold Mask Prediction and Occlusion-aware Convolution for Foreground Occlusions in Light FieldsabstractThe performance of existing de-occlusion methods is limited mainly due to inaccurate foreground mask prediction, interference from occlusion information in feature extractors, and insufficient utilization of sub-pixel information between views. Therefore, in this paper, we propose an efficient light field image de-occlusion method to improve the performance of occlusion localization and occlusion removal. First, we design an adaptive threshold mask prediction branch, which fully utilizes the spatial and angular information of light field images and incorporates mask binarization into the network for joint optimization. Then, we propose Occlusion-aware Convolution, which can more efficiently extract the joint spatial-angular features of the occluded light field image. Additionally, we design a sub-pixel complement strategy. This strategy fully utilizes the sub-pixel information between the views and supplements it into the view to be deoccluded. Experimental results demonstrate that our method achieves superior performance on both real-world and synthetic light field datasets. Jieyu Chen, Ping An 0001, Xinpeng Huang, Chao Yang 0021 |
VCIP | 4 |
| 2024 | Low-Rate Feature Compression for Humans and Machines with Dual Aggregation AttentionabstractThe Collaborative Intelligence (CI) framework offers innovative approaches for deploying Deep Neural Networks (DNNs). However, the limitations of communication resources require minimizing the transmission of bits between the edge and cloud devices to meet the requirements of both machine recognition and human perception. Previous research has demonstrated that the transmission of intermediate layer features of the vision backbone can achieve superior performance in machine vision tasks at very low bit rates without consuming additional bits. Nonetheless, the reduced bit rate poses challenges for image reconstruction. We propose a CI framework that compresses the intermediate features of the Swin Transformer and utilizes a Feature Recovery Module (FRM) to restore crucial information for image reconstruction, thereby satisfying both machine and human visual tasks at low bit rates. Additionally, we introduce a Residual Dual-attention Aggregation Block (RDAB) that exploits both local and global information for effective compression and reconstruction. We conducted experiments on the CUB_200_2011 dataset. The results demonstrate that the proposed method delivers superior performance at low-rate scenarios. Ruixi Ma, Ping An 0001, Shipei Wang, Xinpeng Huang, Chao Yang 0021 |
VCIP | 5 |
| 2024 | Content adaptive spatial-temporal rescaling for video coding optimization
Chao Yang 0021, Siqian Qin, Ping An 0001, Xinpeng Huang, Liquan Shen |
Expert Syst. Appl. | 1 |
| 2024 | Learning-based CU partition prediction for fast panoramic video intra coding
Chao Yang 0021, Ping An 0001, Xinpeng Huang, Liquan Shen |
Expert Syst. Appl. | 2 |
| 2024 | STSIC: Swin-transformer-based scalable image coding for human and machine
Shipei Wang, Ping An 0001, Chao Yang 0021, Kunqiang Huang, Xinpeng Huang |
J. Vis. Commun. Image Represent. | 3 |
| 2024 | Scalable image coding with enhancement features for human and machine
Ping An 0001, Chao Yang 0021, Xinpeng Huang |
Multim. Syst. | 3 |
| 2024 | 360° video quality assessment based on saliency-guided viewport extraction
Fanxi Yang, Chao Yang 0021, Ping An 0001, Xinpeng Huang |
Multim. Syst. | 2 |
| 2024 | Content adaptive downsampling for low bitrate video coding
Siqain Qin, Chao Yang 0021, Ping An 0001 |
Multim. Tools Appl. | 2 |
| 2023 | Intermediate deep feature coding for human-machine vision collaboration
Weiqian Wang, Ping An 0001, Xinpeng Huang, Kunqiang Huang, Chao Yang 0021 |
J. Vis. Commun. Image Represent. | 5 |
| 2022 | An Online SVM Based VVC Intra Fast Partition Algorithm With Pre-Scene-cut DetectionabstractThe new generation of video coding standard, Versatile Video Coding (H.266/VVC), brings tremendous computational complexity by incorporating the quad-tree with nested multi-type tree (QTMT) partition structure. We propose an adaptive low loss fast algorithm to tackle this disadvantage by using the online Support Vector Machine (SVM) classifier. Firstly, we perform a pre-scene-cut detection before encoding the whole sequence to split it into several scenes, which divide frames into training-frame and predicting-frame. Then, the training-frame is used to construct the data set for SVM parameters training. Specifically, we extract partition-related features, i.e., gradient, entropy, and difference of neighbor area depth to train the SVM classifier. Lastly, the partition decision in predicting-frame is accelerated by the SVM classifier model in the same scene with the training-frame. Besides, we control our algorithm to maintain a low Bjontegaard Delta Bit Rate (BDBR) index by applying the SVM classifiers in the most suitable size 32x32. The experimental results show that our algorithm achieves about 15.76% encoding time saving on average with a negligible quality loss-0.23% BDBR increase under all-intra configuration. Chao Shu, Chao Yang 0021, Ping An 0001 |
ISCAS | 2 |
| 2022 | Unsupervised blind image quality assessment based on joint structure and natural scene statistics features
Qinglin He, Chao Yang 0021, Fanxi Yang, Ping An 0001 |
J. Vis. Commun. Image Represent. | 2 |
| 2022 | Light field occlusion removal network via foreground location and background recovery
Shiao Zhang, Ping An 0001, Xinpeng Huang, Chao Yang 0021 |
Signal Process. Image Commun. | 5 |
| 2022 | Objective Quality Assessment of Lenslet Light Field Image Based on Focus StackabstractThe large amount of complex scene information recorded by light field imaging has the potential for immersive media applications. Compression and reconstruction algorithms are crucial for the transmission, storage, and display of such massive data. Most of the existing quality evaluation indexes do not effectively account for light field characteristics. To accurately evaluate the distortions caused by compression and reconstruction algorithms, it is necessary to construct an image evaluation index that reflects the angular-spatial characteristics of the light field. This work proposes a full-reference light field image quality evaluation index that attempts to extract less information from the focus stack to accurately evaluate the entire light field quality. The proposed framework includes three specific steps. First, we construct a key refocused image extraction framework by the maximal spatial information contrast and the minimal angular information variation. Specifically, the gradient and phase congruency operators are used in the extraction framework. Second, a novel light field quality evaluation index is built based on the angular-spatial characteristics of the key refocused images. In detail, the features used in the key refocused image extraction framework and the chrominance feature are combined to construct the union feature. Third, the similarity of the union feature is pooled by the relevant visual saliency map to obtain the predicted score. Finally, the overall quality of the light field is measured by applying the proposed index to the key refocused images. The high efficiency and precision of the proposed method are shown by extensive comparison experiments. Chunli Meng, Ping An 0001, Xinpeng Huang, Chao Yang 0021, Liquan Shen |
IEEE Trans. Multim. | 4 |
| 2021 | Multi-Models Fusion for Light Field Angular Super-ResolutionabstractLight field (LF) imaging has received increasing attention due to its richer interpretation of the scene. However, an inherent spatial-angular trade-off exists in LF that prevents LF from practical applications. Consequently, how to break such a trade-off has become one of the main challenges in sparsely sampled LF reconstruction. LF super-resolution (SR) can provide an opportunity to solve this issue, but most methods exploit only one form of LF, thereby leading to much loss of information. We believe that different LF forms can compensate each other to obtain higher gains via fusion strategy. In this paper, therefore, we propose a multi-models fusion for LF SR in angular domain. Cascading models which are trained by different LF forms can fully exploit rich LF information. Experimental results demonstrate that our method is effective and achieves a comparable result against state-of-the-art techniques. Fengyin Cao, Ping An 0001, Xinpeng Huang, Chao Yang 0021, Qiang Wu 0001 |
ICASSP | 4 |
| 2021 | Blind Image Quality Assessment Based on Multi-scale KLTabstractBlind image quality assessment (BIQA) plays an important role in image services as independent of the reference image. Herein, the perceptual relevant feature design is the core of BIQA methods, but their performance is still not satisfied at present. In this work, we propose an unsupervised feature extraction approach for BIQA based on Karhunen-Loéve transform (KLT). Specifically, a normalization operation is firstly applied to the test image by calculating its mean subtracted contrast normalized (MSCN) coefficient. Then, KLT is employed as a data-driven feature extraction approach to extract image structural features, wherein kernels with different sizes are utilized to perform multi-scale analysis. Finally, generalized Gaussian distribution (GGD) is employed to model the KLT coefficients distribution in different spectral components as quality relevant features. Extensive experiments conducted on four widely utilized IQA databases have demonstrated that the proposed Multi-scale KLT (MsKLT) BIQA metric compares favorably with existing BIQA methods in terms of high accordance with human subjective scores on both common and uncommon distortion types. Chao Yang 0021, Xinfeng Zhang 0001, Ping An 0001, Liquan Shen, C.-C. Jay Kuo |
IEEE Trans. Multim. | 1 |
| 2020 | Post-processing for intra coding through perceptual adversarial learning and progressive refinement
Zhipeng Jin, Ping An 0001, Chao Yang 0021, Liquan Shen |
Neurocomputing | 3 |
| 2020 | Light Field Compression Using Global Multiplane Representation and Two-Step PredictionabstractDue to its spatio-angular structure, light field image allows for a wealth of post-processing techniques like digital refocusing and depth estimation. In order to compress the data of the two domains, the current proposal intends to embed the disparity-based view synthesis method into the decoder. However, predicting each view separately or in local groups means bringing more computational burden to the decoder and destroying the light field structure. Since disparity contains the relationship between all light rays in the light field, the proposed solution is to predict a disparity-based global representation as the first step. In the second step, all the views can be predicted easily based on this representation. In this letter, we use the recently proposed multiplane as the form of this global representation. The experimental results show the effectiveness of the proposed solution, and the better RD performance compared to other schemes especially under low bitrates. Ping An 0001, Xinpeng Huang, Chao Yang 0021, Deyang Liu, Qiang Wu 0001 |
IEEE Signal Process. Lett. | 4 |
| 2020 | Full Reference Light Field Image Quality Evaluation Based on Angular-Spatial CharacteristicabstractThe quality evaluation is an indispensable link in light field (LF) image processing. Most of existing LF objective evaluation indexes do not make effective use of the angular characteristic of LF, so the evaluation results are unsatisfactory. In this letter, the quality evaluation of LF image is constructed based on human visual system (HVS) and LF angular-spatial characteristics. Based on the fact that HVS has different sensitivity to different parallaxes, we assume that LF image quality perceived by human eyes has the optimal parallax range. A dual-fan filter is used to constrain the parallax range. Then, the overall quality of the LF is represented by combining the spatial and angular quality, which performed from the central sub-aperture image and the focus stack, respectively. In addition, because the difference of Gaussian (DoG) operator can simulate the process of extracting texture structure by human eyes. The structural similarity of DoG texture feature is utilized in the spatial quality evaluation. Extensive comparison experiments show that the proposed method is more consistent with the characteristics of LF. Chunli Meng, Ping An 0001, Xinpeng Huang, Chao Yang 0021, Deyang Liu |
IEEE Signal Process. Lett. | 4 |
| 2020 | Image Coding With Data-Driven Transforms: Methodology, Performance and PotentialabstractImage compression has always been an important topic in the last decades due to the explosive increase of images. The popular image compression formats are based on different transforms which convert images from the spatial domain into compact frequency domain to remove the spatial correlation. In this paper, we focus on the exploration of data-driven transform, Karhunen-Loéve transform (KLT), the kernels of which are derived from specific images via Principal Component Analysis (PCA), and design a high efficient KLT based image compression algorithm with variable transform sizes. To explore the optimal compression performance, the multiple transform sizes and categories are utilized and determined adaptively according to their rate-distortion (RD) costs. Moreover, comprehensive analyses on the transform coefficients are provided and a band-adaptive quantization scheme is proposed based on the coefficient RD performance. Extensive experiments are performed on several class-specific images as well as general images, and the proposed method achieves significant coding gain over the popular image compression standards including JPEG, JPEG 2000, and the state-of-the-art dictionary learning based methods. Xinfeng Zhang 0001, Chao Yang 0021, Shan Liu 0001, Haitao Yang 0001, Ioannis Katsavounidis, Shawmin Lei, C.-C. Jay Kuo |
IEEE Trans. Image Process. | 2 |
| 2020 | Satisfied-User-Ratio Modeling for Compressed VideoabstractWith explosive increase of internet video services, perceptual modeling for video quality has attracted more attentions to provide high quality-of-experience (QoE) for end-users subject to bandwidth constraints, especially for compressed video quality. In this paper, a novel perceptual model for satisfied-user-ratio (SUR) on compressed video quality is proposed by exploiting compressed video bitrate changes and spatial-temporal statistical characteristics extracted from both uncompressed original video and reference video. In the proposed method, an efficient video feature set is explored and established to model SUR curves against bitrate variations by leveraging the Gaussian Processes Regression (GPR) framework. In particular, the proposed model is based on the recently released large-scale video quality dataset, VideoSet, and takes both spatial and temporal masking effects into consideration. To make it more practical, we further optimize the proposed method from three aspects including feature source simplification, computation complexity reduction and video codec adaption. Based on experimental results on VideoSet, the proposed method can accurately model SUR curves for various video contents and predict their required bitrates at given SUR values. Subjective experiments are conducted to further verify the generalization ability of the proposed SUR model. Xinfeng Zhang 0001, Chao Yang 0021, Haiqiang Wang, Wei Xu 0001, C.-C. Jay Kuo |
IEEE Trans. Image Process. | 2 |
| 2019 | Objective Quality Assessment for Light Field Based on Refocus Characteristic
Chunli Meng, Ping An 0001, Xinpeng Huang, Chao Yang 0021 |
ICIG (3) | 4 |
| 2018 | Quality Enhancement for Intra Frame Coding Via Cnns: An Adversarial ApproachabstractLossy compression is an indispensable technique in image/video processing, due to its highly desirable ability of reducing the huge data volume. However, lossy compression introduces complex compression artifacts. To reduce these artifacts, post-processing techniques have been extensively studied. In this paper, we propose a novel post-processing technique using multi-level progressive refinement network via an adversarial training approach, called MPRGAN, for artifacts reduction and coding efficiency improvement in intra frame coding. Furthermore, our network generates multi-level residues in one feed-forward pass through the progressive reconstruction. This coarse-to-fine work fashion, which makes our network have high flexibility, can make trade-off between enhanced quality and computational complexity. Thereby facilitates the resource-aware applications. Extensive evaluations on benchmark datasets verify the superiority of our proposed MPRGAN model over the latest state-of-the-art methods with fast deployment running speed. Zhipeng Jin, Ping An 0001, Chao Yang 0021, Liquan Shen |
ICASSP | 3 |
| 2018 | Analysis and Prediction of JND-Based Video Quality ModelabstractThe just-noticeable-difference (JND) visual perception property has received much attention in characterizing human subjective viewing experience of compressed video. In this work, we quantity the JND-based video quality assessment model using the satisfied user ratio (SUR) curve, and show that the SUR model can be greatly simplified since the JND points of multiple subjects for the same content in the VideoSet can be well modeled by the normal distribution. Then, we design an SUR prediction method with video quality degradation features and masking features and use them to predict the first, second and the third JND points and their corresponding SUR curves. Finally, we verify the performance of the proposed SUR prediction method with different configurations on the VideoSet. The experimental results demonstrate that the proposed SUR prediction method achieves good performance in various resolutions with the mean absolute error (MAE) of the SUR smaller than 0.05 on average. Haiqiang Wang, Xinfeng Zhang 0001, Chao Yang 0021, C.-C. Jay Kuo |
PCS | 3 |
| 2018 | Scalable coding of 3D holoscopic image by using a sparse interlaced view image set and disparity map
Deyang Liu, Ping An 0001, Chao Yang 0021, Liquan Shen, Kai Li 0016 |
Multim. Tools Appl. | 4 |
| 2017 | Coding of 3D holoscopic image by using spatial correlation of rendered view imagesabstractHoloscopic imaging is a prospective acquisition and display solution for providing natural and fatigue-free 3D visualization. However, large amount of data is required to represent the 3D holoscopic content. Therefore, efficient coding schemes for this particular type of image are needed. In this paper, an effective coding scheme is proposed by exploring the spatial correlation among the view images with different perspectives rendered from 3D holoscopic image. We utilize the interlaced view image to descript such spatial correlation. A linear prediction method is used on the interlaced view image instead of the original holoscopic image directly. Experimental results show that the proposed coding scheme performs better than HEVC intra standard and screen content coding extension of HEVC with around 2.41dB and 0.42 dB average quality improvement respectively. Deyang Liu, Ping An 0001, Chao Yang 0021, Liquan Shen |
ICASSP | 3 |
| 2017 | CNN oriented fast QTBT partition algorithm for JVET intra codingabstractIn this paper, a novel fast coding unit depth decision algorithm based on convolution neural network is presented for JVET future video coding. JVET employs quad-tree plus binary-tree (QTBT) block partitioning structure, which can support much more flexibility for coding units partition shapes, and improve the coding performance significantly than the HEVC standard. However, the flexible partitioning structure also introduces a tremendous computation complexity. To address this issue, we model the QTBT partition depth range as a multi-class classification problem, and try to predict the depth range of 32×32 block directly, rather than to judge split or not at each depth level. To the best of our knowledge, it is the first framework to formulate the QTBT partition range as a multi classification task, and optimized by an end-to-end learning model. For training optimization, we design an objective function consists of class penalty term and L2 HingeLoss function, which leverage the characteristics of category settings, can further boost the classification accuracy. Experimental results demonstrate the effectiveness of our proposed method, which can achieve 42.80% complexity reduction with only 0.65% Bjontegaard Delta bitrate (BD-rate) increase. Zhipeng Jin, Ping An 0001, Liquan Shen, Chao Yang 0021 |
VCIP | 4 |
| 2017 | Bit allocation for 3D video coding based on lagrangian multiplier adjustment
Chao Yang 0021, Ping An 0001, Deyang Liu, Liquan Shen, Kai Li 0016 |
Signal Process. Image Commun. | 1 |
| 2016 | Depth map coding based on virtual view qualityabstractMulti-view video plus depth (MVD) is a 3D video representation. In MVD, the depth map provides the scene distance information and is used to render the virtual view through Depth Image Based Rendering (DIBR) technique. The depth map coding error will induce distortion in the rendered virtual views. This paper proposes a mathematic model that can estimate the synthesized virtual view distortion induced by depth map compression, and the model is employed to the rate distortion optimization (RDO) in the depth map coding. Based on the rendered virtual view quality, a Lagrangian optimization adjustment scheme at Coding Unit (CU) level is proposed to improve the depth map encoding efficiency. Experimental results demonstrate that the proposed method can improve the BD-PSNR of virtual view for 0.62 dB, and the encoding complexity reduces compared with the view synthesis optimization (VSO) technique in the 3D-HEVC Test Model (HTM). Chao Yang 0021, Ping An 0001, Deyang Liu, Liquan Shen |
ICASSP | 1 |
| 2016 | 3D holoscopic image coding scheme using HEVC with Gaussian process regression
Deyang Liu, Ping An 0001, Chao Yang 0021, Liquan Shen |
Signal Process. Image Commun. | 4 |
| 2016 | Fast depth map coding based on virtual view quality
Chao Yang 0021, Ping An 0001, Liquan Shen, Nina Feng |
Signal Process. Image Commun. | 1 |
| 2015 | Virtual view distortion estimation for depth map codingabstractMulti-view video plus depth (MVD) format is a three-dimensional (3D) video representation. The depth map in MVD provides the scene geometry information and is used to render the virtual view through Depth Image Based Rendering (DIBR). In this paper, a virtual view distortion estimation function based on the characteristics of both texture image and depth map is proposed which can estimate virtual view distortion induced by depth map compression accurately, and the function is implemented to the Rate Distortion Optimization (RDO) in the depth map coding. Compared with the View Synthesis Optimization (VSO) in 3D-HEVC Test Model (HTM) reference software, the experimental results demonstrate that the proposed method can improve the BD-PSNR of virtual view for 0.26 dB on average, and the encoding time has reduced for 31% on average due to the low complexity of the proposed function. Chao Yang 0021, Ping An 0001, Deyang Liu, Liquan Shen |
VCIP | 1 |