VLDB 2026 Research / reviewers in the wild / expert
Ping An 0001
dblp:94/4635-1
· DBLP profile ↗
118ranked-venue papers
0as first author
52since 2021 · last 2026
0000-0002-4995-728XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 104 · 45 since 2021Artificial intelligence and machine learning · 9 · 6 since 2021Human-computer interaction and ubiquitous computing · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3Systems, architecture and hardware · 2 · 1 since 2021Computer networks · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | VIQ-360: A New Viewport-Based Omnidirectional Image Quality Assessment Database
Xiangxu Yu, Chao Yang 0021, Xinpeng Huang, Ping An 0001 |
QoMEX | 5 |
| 2026 | A Subjective Quality Database for Human-AI Co-Created Images
Xiangxu Yu, Jiyan Tong, Chao Yang 0021, Xinpeng Huang, Ping An 0001 |
QoMEX | 7 |
| 2026 | Blind Quality Assessment of Enhanced Low Light Images via Implicit Enhancement Gap Perception
Xiangxu Yu, Zichen Ren, Xinran Gu, Shaoxuan Ding, Ping An 0001 |
QoMEX | 7 |
| 2026 | DMGNet: Discriminative multi-view geometry learning with hybrid-domain enhancement for light field occlusion removal
Jieyu Chen, Ping An 0001, Xinpeng Huang, Chao Yang 0021, Ce Zhu, Sanghoon Lee 0001 |
Expert Syst. Appl. | 2 |
| 2026 | A novel EPI-guided network with progressive fusion for light field reconstruction
Baoshuai Wang, Xinpeng Huang, Ping An 0001 |
Expert Syst. Appl. | 4 |
| 2026 | Low-light light field image enhancement based on illumination-guided implicit gradient representation
Deyang Liu, Xiaofei Zhou 0003, Ping An 0001, Caifeng Shan, Hongbin Zha |
Neurocomputing | 4 |
| 2026 | Enhancing joint human-machine image compression via Chebyshev space modulation
Zhicheng Ma, Ping An 0001, Shipei Wang, Chao Yang 0021, Xinpeng Huang |
Multim. Syst. | 2 |
| 2026 | Generic feature extraction and compression for human and machine-oriented vision
Kunqiang Huang, Ping An 0001, Chao Yang 0021, Shipei Wang, Xinpeng Huang, Liquan Shen |
Signal Process. | 2 |
| 2026 | Viewport-Patch Extraction Enhanced 360$^\circ$ Video Quality AssessmentabstractWith the rising adoption of 360$^\circ$video in virtual reality (VR) applications, assessing its perceptual quality remains a challenge due to projection-induced distortions in equirectangular projection (ERP) formats. Traditional sliding-window cropping methods often distort high-latitude content and fail to reflect the actual viewing experience. To address this, we propose a novel viewport patch-based video quality assessment (VQA) method. By sampling view directions on the sphere and applying gnomonic projection, our method extracts undistorted and perceptually consistent viewport patches that preserve both spatial fidelity and full-frame coverage. We further design a two-stream network that jointly models high-frequency distortion and residual information over time, enhanced by squeeze-and-excitation (SE) attention to capture spatial-temporal features. Experiments and analysis show that our method significantly improves the accuracy and reliability of 360$^\circ$VQA, achieving PLCC/SROCC values of 0.9603/0.9628 on the VQA-ODV dataset and 0.9585/0.9400 on the BIT360 dataset, with only 0.22M parameters. Code is available athttps://github.com/yeonhw/VP-VQA. Chao Yang 0021, Ping An 0001, Xinpeng Huang |
IEEE Signal Process. Lett. | 3 |
| 2026 | Low-Bitrate Light Field Video Compression Through Key Sequences Encoding and Joint Reconstruction NetworkabstractLight field (LF) videos contain rich spatial, angular, and temporal information, resulting in immense data volumes and posing significant challenges for low-bitrate compression. Existing LF video compression methods focus on modifying the structure of traditional video codecs to encode all LF views, but they are insufficient to achieve low-bitrate compression of LF video. To address these limitations, we propose a low-bitrate LF video compression framework that exploits spatial-angular-temporal correlations through sparse coding and joint reconstruction. On the encoding side, we introduce a content-adaptive prediction structure for sparse key view sequences selection. This structure is adapted to LF video content, leveraging the most similar view as a reference to enhance prediction accuracy and significantly reduce bitrate. On the decoding side, we observe that pixels missing in the current view are often captured in adjacent angular and/or temporal views. As a result, we develop a spatial-angular-temporal based joint reconstruction network that integrates cues across the different domains. This approach supplements missing texture details near occlusion areas and reconstructs high-quality non-key views. Experimental results demonstrate the efficiency of our framework, achieving an average gain of about 60 % in terms of bitrate savings and 2 dB in terms of reconstruction quality compared to the state-of-the-art methods. Xinpeng Huang, Chao Yang 0021, Mounir Kaaniche, Qiuwen Zhang, Ping An 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2026 | Capture More, Synthesize Better: Video Frame Interpolation With Larger Receptive Field and Structural PriorsabstractA large receptive field is crucial for the video frame interpolation (VFI) task. Existing video frame interpolation methods struggle with large motions due to their limited receptive fields. However, simply expanding the receptive field brings two challenges: a substantial computational burden and potential loss of texture details. In this paper, we first propose a novel spatial-temporal global window self-attention mechanism with an enlarged receptive field to enhance motion capture. Furthermore, to reduce the computational complexity introduced by the global window, we design a simple and effective separable fence window decomposition. Meanwhile, to better synthesize high-quality intermediate frames, we propose two complementary frame synthesis strategies. First, from the perspective of receptive field design, we introduce a progressive receptive field focusing module, enabling a smooth transition from global motion modeling to local detail preservation. Second, based on the VFI-specific property and the high structural similarity shared by the adjacent frames, we propose a structure-aware synthesis strategy, which incorporates structural priors to guide the generation of fine details. Subjective and objective experimental results demonstrate that our method effectively captures large motions while synthesizing texture details, outperforming state-of-the-art techniques on various datasets. Baojun Zhou, Xinpeng Huang, Jieyu Chen, Mounir Kaaniche, Ping An 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2026 | Volume Feature Aware View-Epipolar Transformers for Generalizable NeRFabstractGeneralizable NeRF synthesizes novel views of unseen scenes without per-scene training. The view-epipolar transformer has become popular in this field for its ability to produce high-quality views. Existing methods with this architecture rely on the assumption that texture consistency across views can identify object surfaces, with such identification crucial for determining where to reconstruct texture. However, this assumption is not always valid, as different surface positions may share similar texture features, creating ambiguity in surface identification. To handle this ambiguity, this paper introduces 3D volume features into the view-epipolar transformer. These features contain geometric information, which will be a supplement to texture features. By incorporating both texture and geometric cues in consistency measurement, our method mitigates the ambiguity in surface detection. This leads to more accurate surfaces and thus better novel view synthesis. Additionally, we propose a decoupled decoder where volume and texture features are used for density and color prediction respectively. In this way, the two properties can be better predicted without mutual interference. Experiments show improved results over existing transformer-based methods on both real-world and synthetic datasets. Ping An 0001, Xinpeng Huang, Qiang Wu 0001 |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2026 | Learning a Domain-Specialized Network for Light Field Spatial-Angular Super-ResolutionabstractLight field (LF) imaging is inherently constrained by the trade-off between spatial resolution and angular sampling density. To overcome this obstacle, spatial-angular super-resolution (SR) methods have been developed to achieve concurrent enhancement in both dimensions. Traditional spatial-angular SR methods treat spatial and angular SR as separate tasks, resulting in parameter redundancy and error accumulation. While recent end-to-end approaches attempt joint processing, their uniform treatment of these distinct problems overlooks critical domain-specific requirements. To address these challenges, we propose a domain-specialized framework that deploys stage-tailored strategies to satisfy domain-specific demands. Specifically, in the angular SR stage, we introduce a cross-view consistency modulation module that enhances inter-view coherence through long-range dependency modeling of angular features. In the spatial SR stage, we propose a detail-aware state space model to reconstruct fine-grained detail. Finally, we develop a cross-domain integration module that explores spatial-angular correlations by fusing multi-representational features from both domains to foster synergistic optimization. Experimental results on public LF datasets demonstrate substantial improvements over state-of-the-art methods in both qualitative and quantitative comparisons, with approximately 50% fewer model parameters compared to competing methods. Xinpeng Huang, Deyang Liu, Ping An 0001, Sanghoon Lee 0001 |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2025 | Domain Crossover Non-Rigid Registration for 3D Human MeshesabstractNon-rigid registration is essential for reconstructing dynamic and incomplete 3D human meshes, yet traditional methods often fail to achieve robust alignment in the sequence of high-motion deformations and missing geometry. We propose a domain crossover non-rigid registration (DCNRR) framework that addresses these challenges by effectively transferring informative features from 2D image space into the 3D mesh domain of three key stages: multi-view projection, hierarchical non-rigid registration, and topology-consistent completion. In the first stage, multi-view projections are used to extract 2D joint locations and deep features, which guide deformation in the 3D space. In the second stage, hierarchical joint priors and deep features collaboratively guide mesh alignment, enabling more accurate deformation in distal regions and complex poses. In the final stage, we apply a diffusion-based completion process in UV coordinates to reconstruct incomplete surface normals and refine missing mesh areas with topological consistency. Our approach achieves highly detailed and perceptually accurate mesh deformation. To validate our approach, we evaluate performance on a newly constructed dynamic human motion (DHM) dataset, as well as public datasets. Our method demonstrates state-of-the-art results in both geometric accuracy and stability, showing particular robustness in dynamic and incomplete mesh sequences. Kyungjune Lee, Seongjean Kim, Hoseok Tong, Hyucksang Lee, Seongmin Lee 0002, Weisi Lin, Ping An 0001, Sanghoon Lee 0001 |
ACM Multimedia | 7 |
| 2025 | Layered and scalable image coding with semantic features for human and machine
Jiao Wei, Ping An 0001, Shipei Wang, Kunqiang Huang, Chao Yang 0021, Xinpeng Huang |
Eng. Appl. Artif. Intell. | 2 |
| 2025 | Spherical rotation for high efficiency ERP 360-degree video coding
Tianyu Hong, Guowei Teng, Ping An 0001, Liquan Shen |
Multim. Syst. | 3 |
| 2025 | Disparity Enhancement-Based Light Field Angular Super-ResolutionabstractThe depth-dependent light field (LF) reconstruction is a prevalent solution for large disparity LFs, which first estimates disparity maps and subsequently interpolates the content of target views by warping input views. However, replication errors often occur in edge regions of objects owing to inappropriate sampling positions caused by occlusion during disparity-based warping. Thus, we propose a disparity enhancement network that utilizes morphological filtering to address this distortion, which can adaptively modify disparity values in edge regions to obtain proper sampling positions. In addition, we develop an effective detail recovery network to mitigate interpolation errors introduced by inaccurate disparity estimation or warping operations. Experiments demonstrate that our approach significantly surpasses current state-of-the-art methods in large disparity LFs. Dongjun Cai, Xinpeng Huang, Ping An 0001 |
IEEE Signal Process. Lett. | 4 |
| 2025 | DiffHSR: Unleashing Diffusion Priors in Hyperspectral Image Super-ResolutionabstractHyperspectral images provide rich spectral information and have been widely applied in numerous computer vision tasks. However, their low spatial resolution often limits their use in applications such as image segmentation and recognition. In previous works, generating high-resolution hyperspectral (HR-HS) images required the use of low-resolution hyperspectral (LR-HS) images and high-resolution RGB (HR-RGB) images as priors, which increases the cost of data collection and may lead to measurement and calibration errors in practical applications. Although the currently popular CNN-based single hyperspectral image super-resolution (single HS-SR) methods have improved performance, they are not flexible enough to process images with different degradation. From a visual perspective, the generated super-resolution images exhibit a significant smudging effect due to the loss of information. Leveraging multi-modal techniques and generative prior, we propose DiffHSR that marks a significant leap in LR-HS images super-restoration without HR-RGB. Additionally, we have established a connection between hyperspectral images and the RGB image-based generative model tasks using low-cost data and fine-tuning approaches, which creates a novel paradigm. Comprehensive experiments have demonstrated that our proposed method achieves strong visual performance and competitive results in term of quantitative metrics and perceptive quality. Yumeng Xie, Ping An 0001 |
IEEE Signal Process. Lett. | 3 |
| 2025 | L3FMamba: Low-Light Light Field Image Enhancement With Prior-Injected State Space ModelsabstractIn this paper, we address the problem of low-light light field (LF) image enhancement, where spatial details and angular coherence are severely degraded due to noise and insufficient illumination. Existing methods often rely on local aggregation or naive view stacking, which fail to capture global illumination and long-range spatial-angular correlations. To overcome these limitations, we propose L3FMamba, a lightweight enhancement method that integrates Retinex and Atmospheric Scattering models with dark, bright, and average channel priors for robust illumination decomposition. Moreover, we incorporate a state space model to capture non-local spatial-angular dependencies, enabling effective propagation of global context across views. By combining physics-inspired priors with structured modeling, L3FMamba achieves accurate illumination correction and fine-detail preservation with minimal parameters. Experiments show that L3FMamba outperforms the state-of-the-art in quality. Deyang Liu, Shizheng Li, Zeyu Xiao 0002, Ping An 0001, Caifeng Shan |
IEEE Signal Process. Lett. | 4 |
| 2025 | Deep Sparse-to-Dense Inbetweening for Multi-View Light FieldsabstractLight field (LF) imaging, which captures both intensity and directional information of light rays, extends the capabilities of traditional imaging techniques. In this paper, we introduce a task in the field of LF imaging, sparse-to-dense inbetweening, which focuses on generating dense novel views from sparse multi-view LFs. By synthesizing intermediate views from sparse inputs, this task enhances LF view synthesis through filling in interperspective gaps within an expanded field of view and increasing data robustness by leveraging complementary information between light rays from different perspectives, which are limited by non-robust single-view synthesis and the inability to handle sparse inputs effectively. To address these challenges, we construct a high-quality multi-view LF dataset, consisting of 60 indoor scenes and 59 outdoor scenes. Building upon this dataset, we propose a baseline method. Specifically, we introduce an adaptive alignment module to dynamically align information by capturing relative displacements. Next, we explore angular consistency and hierarchical information using a multi-level feature decoupling module. Finally, a multi-level feature refinement module is applied to enhance features and facilitate reconstruction. Additionally, we introduce a universally applicable artifact-aware loss function to effectively suppress visual artifacts. Experimental results demonstrate that our method outperforms existing approaches, establishing a benchmark for sparse-to-dense inbetweening. The code is available at https://github.com/Starmao1/MutiLF. Zeyu Xiao 0002, Ping An 0001, Deyang Liu, Caifeng Shan |
IEEE Trans. Image Process. | 3 |
| 2025 | Mask-Aware Light Field De-Occlusion With Gated Feature Aggregation and Texture-Semantic AttentionabstractA light field image records rich information of a scene from multiple views, thereby providing complementary information for occlusion removal. However, current occlusion removal methods have several issues: 1) inefficient exploitation of spatial and angular complementary information among views; 2) indistinguishable treatment of pixels from foreground occlusion and background; and 3) insufficient exploration of spatial detail supplementation. Therefore, in this article, we propose a mask-aware de-occlusion network (MANet). Specifically, MANet is a joint training network that integrates the occlusion mask predictor (OMP) and the occlusion remover (OR). First, OMP is proposed to provide the location of occluded regions for OR, as the occlusion removal task is ill-posed without occluded region localization. In OR, we introduce gated spatial-angular feature aggregation, which uses a soft gating mechanism to focus on spatial-angular interaction features in non-occluded regions, extracting effective aggregated features specific to the de-occlusion. Then, we design a complementary strategy to fully utilize spatial-angular information among views. Finally, we propose texture-semantic attention to improve the performance of detail generation. Experimental results demonstrate the superiority of MANet, with substantial improvements in both PSNR and SSIM metrics. Moreover, MANet stands out with an efficient parameter count of 2.4 M, making it a promising solution for real-world applications in public safety and security surveillance. Jieyu Chen, Ping An 0001, Xinpeng Huang, Chao Yang 0021, Liquan Shen |
IEEE Trans. Multim. | 2 |
| 2025 | Feature Quality Assessment: A Database and A Lightweight Objective MethodabstractIn the era of Artificial Intelligence, visual data gathered by edge devices could be primarily utilized for machine vision tasks. The prominent coding frameworks accomplish this by extracting and compressing features extracted from input data. As such, the quality of these features is vital, as they reflect the performance of the coding framework. However, much less work has been dedicated to quality assessment on features, impeding the optimization of the coding system. In this work, we pioneer to explore the feature quality assessment by creating a novel database tailored for features, with the quality ground-truth for each feature. Then, we propose a lightweight feature quality assessment method, called Lightweight Feature Quality Assessment (LFQA). We analyze the feature characteristics from the perspective of spatial and channel thoroughly, and the framework of LFQA is designed based on the analysis results. Experimental results demonstrate that LFQA accurately evaluates the quality of features, reaching a notable Spearman Rank-Order Correlation Coefficient of 85.38%, and exhibits competitive performance in improving the performance of video coding for machine system. Furthermore, LFQA has fewer model parameters and faster inference speed, ensuring a wide range of promising applications. Shipei Wang, Ping An 0001, Chao Yang 0021, Gongyang Li, Xinpeng Huang, Shiqi Wang 0001 |
IEEE Trans. Multim. | 2 |
| 2025 | Dual-Guided Video Frame Interpolation With Spatial-Temporal Global AttentionabstractVideo frame interpolation technology improves visual experience with the development of deep learning. However, capturing large motions while synthesizing fine texture details remains a challenging task. Regarding large motion scenarios, some pioneering Transformer-based methods primarily rely on local attention, which does not fully leverage the global receptive field advantage. To address this issue, this paper proposes to further broaden the receptive field of the Transformer to capture more correlations in the video frame interpolation task. Specifically, we propose a global self-attention mechanism in the form of spatial-temporal separation. Regarding texture details, since roughly enlarging the receptive field results in the loss of details, we propose to use large motion information in both feature and pixel spaces as a dual-guided prior to enhance detail synthesis. The separable attention mechanism and the straightforward frame synthesis design significantly enhance the resource efficiency of our model. Extensive experiments show that our method achieves state-of-the-art performance, effectively capturing large motions and preserving texture details. Baojun Zhou, Xinpeng Huang, Gongyang Li, Chao Yang 0021, Liquan Shen, Ping An 0001 |
IEEE Trans. Multim. | 6 |
| 2025 | Towards 360 VR Sickness Mitigation: From Virtual Reality Eye-Tracking to Visual CommunicationabstractMost 360 virtual reality (VR) contents have been developed without considering that users could be affected by VR sickness. Accordingly, users' viewing safety has been steadily highlighted as a critical problem in the VR market. In this study, we investigate a novel VR sickness mitigation framework based on human visual characteristics for the rendered VR content. First, we build a large-scale 360 VR content database termed VRSP360 (VR Sickness and Presence 360) dedicated to the analysis of VR sickness and thoroughly conduct eye-tracking experiments to measure human perception. In the experiment, we observe that the users' gaze distribution is highly center-biased when they experience excessive VR sickness. From this observation, we design a foveated filtering framework that limits high-frequency textures in the peripheral view to mitigate VR sickness. Particularly, given the human visual system's (HVS) non-uniform resolution with respect to the fovea, we also adopt the foveation-based filtering method using the trade-off between sickness mitigation and presence conservation, which reduces any loss in perceptual quality despite the filtering. We further demonstrate that our framework can effectively compress visual information by applying foveated compression. In addition, we develop two metrics (visual texture index and perceptual information index) to measure the effective preservation of user-perceived information despite the filtration of peripheral vision textures by our proposed mitigation method. Through rigorous subjective evaluation on both original content and its VR-sickness-mitigated version, we demonstrate that the proposed framework successfully mitigates VR sickness with a reduction rate of $\sim$∼19% on the proposed dataset. Jeonghaeng Lee, Woojae Kim, Chao Yang 0021, Ping An 0001, Sanghoon Lee 0001 |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2024 | Adaptive Threshold Mask Prediction and Occlusion-aware Convolution for Foreground Occlusions in Light FieldsabstractThe performance of existing de-occlusion methods is limited mainly due to inaccurate foreground mask prediction, interference from occlusion information in feature extractors, and insufficient utilization of sub-pixel information between views. Therefore, in this paper, we propose an efficient light field image de-occlusion method to improve the performance of occlusion localization and occlusion removal. First, we design an adaptive threshold mask prediction branch, which fully utilizes the spatial and angular information of light field images and incorporates mask binarization into the network for joint optimization. Then, we propose Occlusion-aware Convolution, which can more efficiently extract the joint spatial-angular features of the occluded light field image. Additionally, we design a sub-pixel complement strategy. This strategy fully utilizes the sub-pixel information between the views and supplements it into the view to be deoccluded. Experimental results demonstrate that our method achieves superior performance on both real-world and synthetic light field datasets. Jieyu Chen, Ping An 0001, Xinpeng Huang, Chao Yang 0021 |
VCIP | 2 |
| 2024 | Low-Rate Feature Compression for Humans and Machines with Dual Aggregation AttentionabstractThe Collaborative Intelligence (CI) framework offers innovative approaches for deploying Deep Neural Networks (DNNs). However, the limitations of communication resources require minimizing the transmission of bits between the edge and cloud devices to meet the requirements of both machine recognition and human perception. Previous research has demonstrated that the transmission of intermediate layer features of the vision backbone can achieve superior performance in machine vision tasks at very low bit rates without consuming additional bits. Nonetheless, the reduced bit rate poses challenges for image reconstruction. We propose a CI framework that compresses the intermediate features of the Swin Transformer and utilizes a Feature Recovery Module (FRM) to restore crucial information for image reconstruction, thereby satisfying both machine and human visual tasks at low bit rates. Additionally, we introduce a Residual Dual-attention Aggregation Block (RDAB) that exploits both local and global information for effective compression and reconstruction. We conducted experiments on the CUB_200_2011 dataset. The results demonstrate that the proposed method delivers superior performance at low-rate scenarios. Ruixi Ma, Ping An 0001, Shipei Wang, Xinpeng Huang, Chao Yang 0021 |
VCIP | 2 |
| 2024 | Content adaptive spatial-temporal rescaling for video coding optimization
Chao Yang 0021, Siqian Qin, Ping An 0001, Xinpeng Huang, Liquan Shen |
Expert Syst. Appl. | 3 |
| 2024 | Learning-based CU partition prediction for fast panoramic video intra coding
Chao Yang 0021, Ping An 0001, Xinpeng Huang, Liquan Shen |
Expert Syst. Appl. | 3 |
| 2024 | STSIC: Swin-transformer-based scalable image coding for human and machine
Shipei Wang, Ping An 0001, Chao Yang 0021, Kunqiang Huang, Xinpeng Huang |
J. Vis. Commun. Image Represent. | 2 |
| 2024 | Scalable image coding with enhancement features for human and machine
Ping An 0001, Chao Yang 0021, Xinpeng Huang |
Multim. Syst. | 2 |
| 2024 | 360° video quality assessment based on saliency-guided viewport extraction
Fanxi Yang, Chao Yang 0021, Ping An 0001, Xinpeng Huang |
Multim. Syst. | 3 |
| 2024 | Content adaptive downsampling for low bitrate video coding
Siqain Qin, Chao Yang 0021, Ping An 0001 |
Multim. Tools Appl. | 3 |
| 2024 | Light Field Salient Object Detection With Sparse Views via Complementary and Discriminative Interaction Networkabstract4D light field data record the scene from multiple views, thus implicitly providing beneficial depth cue for salient object detection in challenging scenes. Existing light field salient object detection (LF SOD) methods usually use a large number of views to improve the detection accuracy. However, using so many views for LF SOD brings difficulties to its practical applications. Considering that adjacent views in a light field are actually with very similar contents, in this work, we propose defining a more efficient pattern of input views, i. e., key sparse views, and design a network to effectively explore the depth cue from sparse views for LF SOD. Specifically, we firstly introduce a low rank-based statistical analysis to the existing LF SOD datasets, which allows us to conclude a fixed yet universal pattern for our key sparse views, including the number and positions of views. These views maintain the sufficient depth cue, but greatly lower the number of views to be captured and processed, facilitating practical applications. Then, we propose an effective solution with a key Complementary and Discriminative Interaction Module (CDIM) for LF SOD from key sparse views, named CDINet. The CDINet follows a two-stream structure to extract the depth cue from the light field stream (i. e., sparse views) and the appearance cue from the RGB stream (i. e., center view), generating features and initial saliency maps for each stream. The CDIM is tailored for inter-stream interaction of both these features and saliency maps, using the depth cue to complement the missing salient regions in RGB stream and discriminate the background distraction, to enhance the final saliency map further. Extensive experiments on three LF multi-view datasets demonstrate that our CDINet not only outperforms the state-of-the-art 2D methods, but also achieves competitive performance as compared with the state-of-the-art 3D and 4D methods. The code and results of our method are available athttps://github.com/GilbertRC/LFSOD-CDINet. Gongyang Li, Ping An 0001, Zhi Liu 0003, Xinpeng Huang, Qiang Wu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | Light Field Image Quality Assessment Using Natural Scene Statistics and Texture DegradationabstractLight field image (LFI) now is becoming increasingly popular in immersive media applications. Unlike traditional 2D and 3D images, images taken by light field cameras can capture both angular and spatial information. However, the spatial and angular information of LFI is highly inter-twined with varying disparities, which poses a higher challenge to the quality assessment of LFI. To address this issue, this paper proposes a full-reference light field image quality assessment (LFIQA) index that attempts to disentangle the coupling information from macro-pixel image (MacPI) to accurately evaluate the entire LFI quality. The proposed framework can be divided into three steps. Firstly, the LFIs are converted into the MacPIs, and then the spatial and angular feature maps are disentangled by using the spatial, angular and epipolar plane image (EPI) convolutions in the MacPI mode. Secondly, the structural similarity (SSIM) maps are calculated between the disentangled feature maps of the original and distorted LFIs. Furthermore, the quality-aware features of LFIs are extracted on the SSIM maps by utilized local binary patterns (LBP) and natural scene statistics (NSS). Finally, support vector regression (SVR) is utilized to predict the qualities of LFIs. Extensive experiments show that the proposed model outperforms multiple classical and state-of-the-art methods. Jian Ma 0012, Xiaoyin Zhang, Ping An 0001, Guoming Xu |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | meTMQI: multi-task and exposure-prior learning for Tone-Mapped Quality Index
Mingxing Jiang, Liquan Shen, Xiangyu Hu 0003, Min Hu 0010, Ping An 0001, Tao Tian |
Vis. Comput. | 5 |
| 2023 | Learning a Multilevel Cooperative View Reconstruction Network for Light Field Angular Super-ResolutionabstractRecently, many methods have been proposed to improve the angular resolution of sparsely-sampled Light Field (LF). However, the synthesized dense LF inevitably exhibits blurry edges and artifacts. This paper intents to model the global relations of LF views and quality degradation model by learning a multilevel cooperative view reconstruction network to further enhance LF angular Super-Resolution (SR) performance. The proposed LF angular SR network consists of three sub-networks including the Cooperative Angular Transformer Network (CATNet), the Deblurring Network (DBNet), and the Texture Repair Network (TRNet). The CATNet simultaneously captures global features of all LF views and local features within each view, which benefits in characterizing the inherent LF structure. The DBNet models a quality degradation model by estimating blur kernels to reduce the blurry edges and artifacts. The TRNet focuses on restoring fine-scale texture details. Experimental results over various LF datasets including large baseline LF images demonstrate the significant superiority of our method when compared with state-of-the-art ones. Deyang Liu, Xiaofei Zhou 0003, Ping An 0001, Yuming Fang 0001 |
ICME | 4 |
| 2023 | Intermediate deep feature coding for human-machine vision collaboration
Weiqian Wang, Ping An 0001, Xinpeng Huang, Kunqiang Huang, Chao Yang 0021 |
J. Vis. Commun. Image Represent. | 2 |
| 2023 | LA-HDR: Light Adaptive HDR Reconstruction Framework for Single LDR Image Considering Varied Light ConditionsabstractThe high dynamic range (HDR) image recovery from the low dynamic range (LDR) image aims to estimate HDR image by decompressing luminance range and enhancing details of the LDR input. In practical usages, when faced with the over-exposed, the under-exposed or the low-light images, the state-of-art prediction methods lack the capability for ideally handling them. Aiming for this, a light adaptation HDR recovery framework (LA-HDR) is proposed, which includes the multi-images generation for adaptive details amplification in different light ranges, and the following multi-details fusion. To create the multi-images, first, the designed bit-depth enhancement network (EnhanceNet) produces the high bit-depth result with enhanced contrast. This result can be furtherly processed by user-defined denoising method to refrain the low-light noise. Meanwhile, the proposed exposure bias network (EBNet) estimates the global exposure bias of the input for rectifying the mid-range details. With the enhanced result and the exposure bias, the designed transfer functions adaptively create three multi-images containing the enhanced details in different light ranges, and they are fused by the designed multi-images fusion network (FuseNet) for the final HDR prediction. The amplification and fusion scheme ensures robust HDR recovery under different light conditions, eliminating high-light recovery artifacts from previous methods. The proposed fusion masks generation (FMG) and the global feature embedding (GFE) modules inFuseNethelp eliminate the fusion artifacts. Experimental results show that LA-HDR acquires the best average performance under various light conditions, and it receives low influence from the input light conditions among the tested state-of-art HDR recovery methods. Xiangyu Hu 0003, Liquan Shen, Mingxing Jiang, Ping An 0001 |
IEEE Trans. Multim. | 5 |
| 2023 | Multi-Stream Dense View Reconstruction Network for Light Field Image CompressionabstractRecently, many view synthesis-based methods are proposed for high-efficiency light field (LF) image compression. However, most existing methods fail to recover more texture details on occlusion regions, which reduces the compression efficiency. In this paper, we propose a multi-stream dense view reconstruction network to further improve LF image compression performance. In our method, only sparsely-sampled LF views are transmitted and the rest of the views are reconstructed at the decoder side. During the reconstruction process, we firstly constitute a multi-disparity geometry (MDG) structure based on the decoded sparse LF views, which can reflect abundant disparity characteristics. Subsequently, a multi-stream view reconstruction network (MSVRNet) is put forward to reconstruct a high-quality dense LF image, which consists of a multi-scale feature fusion sub-network, a fusion reconstruction sub-network, and a detail refinement sub-network. The multi-scale feature fusion sub-network can implicitly lean abundant multiscale geometric structure features from the constituted MDG structure. The fusion reconstruction sub-network and the detail refinement sub-network are respectively utilized to fuse the learned multiscale geometric features and restore more texture details, especially for occlusion regions. Moreover, 3D convolutional operations are adopted in the whole reconstruction process, which allow information propagation among the learned multiscale geometric features. Comprehensive experimental results demonstrate the effectiveness of the proposed method. The perceptual quality of reconstructed views and application on depth estimation also demonstrate that the proposed method can keep structural consistency of the reconstructed LF image and recover more texture details. Deyang Liu, Yan Huang 0023, Yuming Fang 0001, Yifan Zuo 0001, Ping An 0001 |
IEEE Trans. Multim. | 5 |
| 2022 | An Online SVM Based VVC Intra Fast Partition Algorithm With Pre-Scene-cut DetectionabstractThe new generation of video coding standard, Versatile Video Coding (H.266/VVC), brings tremendous computational complexity by incorporating the quad-tree with nested multi-type tree (QTMT) partition structure. We propose an adaptive low loss fast algorithm to tackle this disadvantage by using the online Support Vector Machine (SVM) classifier. Firstly, we perform a pre-scene-cut detection before encoding the whole sequence to split it into several scenes, which divide frames into training-frame and predicting-frame. Then, the training-frame is used to construct the data set for SVM parameters training. Specifically, we extract partition-related features, i.e., gradient, entropy, and difference of neighbor area depth to train the SVM classifier. Lastly, the partition decision in predicting-frame is accelerated by the SVM classifier model in the same scene with the training-frame. Besides, we control our algorithm to maintain a low Bjontegaard Delta Bit Rate (BDBR) index by applying the SVM classifiers in the most suitable size 32x32. The experimental results show that our algorithm achieves about 15.76% encoding time saving on average with a negligible quality loss-0.23% BDBR increase under all-intra configuration. Chao Shu, Chao Yang 0021, Ping An 0001 |
ISCAS | 3 |
| 2022 | Unsupervised blind image quality assessment based on joint structure and natural scene statistics features
Qinglin He, Chao Yang 0021, Fanxi Yang, Ping An 0001 |
J. Vis. Commun. Image Represent. | 4 |
| 2022 | An improved lossless image compression algorithm based on Huffman coding
Ping An 0001, Xinpeng Huang |
Multim. Tools Appl. | 2 |
| 2022 | A novel deep translated attention hashing for cross-modal retrieval
Ping An 0001, Kai Li 0016 |
Multim. Tools Appl. | 4 |
| 2022 | Energy-driven reference selection for hierarchical light field compression
Xinpeng Huang, Ping An 0001, Deyang Liu |
Signal Process. Image Commun. | 2 |
| 2022 | Light field occlusion removal network via foreground location and background recovery
Shiao Zhang, Ping An 0001, Xinpeng Huang, Chao Yang 0021 |
Signal Process. Image Commun. | 3 |
| 2022 | Low Bitrate Light Field Compression With Geometry and Content ConsistencyabstractLight field imaging can simultaneously record the position and direction information of light rays; thus, digital refocusing and full depth-of-field extension — functions that are inaccessible for conventional images — can be achieved using the structural consistency of light field data. To meet the challenges of limited bandwidth and storage, such vast numbers of light field data must be compressed to a low bitrate. However, current compression solutions ignore the intrinsic consistency of light fields in pursuit of a low bitrate, thereby leading to the loss of light field capabilities. To solve this issue, this work focuses on structural consistency to achieve efficient light field compression with a low bitrate. The proposed light field compression method encodes the sparsely selected sub-aperture images (SAIs) and the disparity maps corresponding to the unselected SAIs. From the perspective of geometry consistency, the consistency of the initially estimated disparity maps is improved by using a color-guided refinement algorithm, thereby reducing the bitrate of the disparity maps. From the perspective of content consistency, the consistency of the SAI-transformed pseudo sequence is improved by the proposed content-similarity-based arrangement algorithm along with a specific prediction structure; thereby, the bitrate of the sparsely selected SAIs is reduced. The experimental results show that the proposed compression method can reduce the total bitrate while preserving good structural consistency. Xinpeng Huang, Ping An 0001, Deyang Liu, Liquan Shen |
IEEE Trans. Multim. | 2 |
| 2022 | Objective Quality Assessment of Lenslet Light Field Image Based on Focus StackabstractThe large amount of complex scene information recorded by light field imaging has the potential for immersive media applications. Compression and reconstruction algorithms are crucial for the transmission, storage, and display of such massive data. Most of the existing quality evaluation indexes do not effectively account for light field characteristics. To accurately evaluate the distortions caused by compression and reconstruction algorithms, it is necessary to construct an image evaluation index that reflects the angular-spatial characteristics of the light field. This work proposes a full-reference light field image quality evaluation index that attempts to extract less information from the focus stack to accurately evaluate the entire light field quality. The proposed framework includes three specific steps. First, we construct a key refocused image extraction framework by the maximal spatial information contrast and the minimal angular information variation. Specifically, the gradient and phase congruency operators are used in the extraction framework. Second, a novel light field quality evaluation index is built based on the angular-spatial characteristics of the key refocused images. In detail, the features used in the key refocused image extraction framework and the chrominance feature are combined to construct the union feature. Third, the similarity of the union feature is pooled by the relevant visual saliency map to obtain the predicted score. Finally, the overall quality of the light field is measured by applying the proposed index to the key refocused images. The high efficiency and precision of the proposed method are shown by extensive comparison experiments. Chunli Meng, Ping An 0001, Xinpeng Huang, Chao Yang 0021, Liquan Shen |
IEEE Trans. Multim. | 2 |
| 2021 | Multi-Models Fusion for Light Field Angular Super-ResolutionabstractLight field (LF) imaging has received increasing attention due to its richer interpretation of the scene. However, an inherent spatial-angular trade-off exists in LF that prevents LF from practical applications. Consequently, how to break such a trade-off has become one of the main challenges in sparsely sampled LF reconstruction. LF super-resolution (SR) can provide an opportunity to solve this issue, but most methods exploit only one form of LF, thereby leading to much loss of information. We believe that different LF forms can compensate each other to obtain higher gains via fusion strategy. In this paper, therefore, we propose a multi-models fusion for LF SR in angular domain. Cascading models which are trained by different LF forms can fully exploit rich LF information. Experimental results demonstrate that our method is effective and achieves a comparable result against state-of-the-art techniques. Fengyin Cao, Ping An 0001, Xinpeng Huang, Chao Yang 0021, Qiang Wu 0001 |
ICASSP | 2 |
| 2021 | An optimized CNN-based quality assessment model for screen content image
Xuhao Jiang, Liquan Shen, Guorui Feng, Liangwei Yu, Ping An 0001 |
Signal Process. Image Commun. | 5 |
| 2021 | Learning from EPI-Volume-Stack for Light Field image angular super-resolution
Deyang Liu, Qiang Wu 0001, Yan Huang 0023, Xinpeng Huang, Ping An 0001 |
Signal Process. Image Commun. | 5 |
| 2021 | Blind Image Quality Assessment Based on Multi-scale KLTabstractBlind image quality assessment (BIQA) plays an important role in image services as independent of the reference image. Herein, the perceptual relevant feature design is the core of BIQA methods, but their performance is still not satisfied at present. In this work, we propose an unsupervised feature extraction approach for BIQA based on Karhunen-Loéve transform (KLT). Specifically, a normalization operation is firstly applied to the test image by calculating its mean subtracted contrast normalized (MSCN) coefficient. Then, KLT is employed as a data-driven feature extraction approach to extract image structural features, wherein kernels with different sizes are utilized to perform multi-scale analysis. Finally, generalized Gaussian distribution (GGD) is employed to model the KLT coefficients distribution in different spectral components as quality relevant features. Extensive experiments conducted on four widely utilized IQA databases have demonstrated that the proposed Multi-scale KLT (MsKLT) BIQA metric compares favorably with existing BIQA methods in terms of high accordance with human subjective scores on both common and uncommon distortion types. Chao Yang 0021, Xinfeng Zhang 0001, Ping An 0001, Liquan Shen, C.-C. Jay Kuo |
IEEE Trans. Multim. | 3 |
| 2021 | Frequency-Dependent Depth Map Enhancement via Iterative Depth-Guided Affine Transformation and Intensity-Guided RefinementabstractRecently, deep convolutional neural network sho-ws significant improvement for intensity-guided depth map enhancement. The most networks focus on either increasing depth or easing features propagation via residual learning and dense connection. However, it has not been explicitly considered yet to mitigate the artifacts caused by the differences of the distributions between the depth map and the corresponding color image, e.g., edge misalignment. In this paper, a novel depth-guided affine transformation is used to filter out the unrelated intensity features, which is further used to refine the depth features. Since the quality of initial depth features is low, the depth-guided intensity features filtering and the intensity-guided depth features refinement are iteratively performed, which progressively promotes effects of such tasks. To make full use of the iterations, all the refined depth features are dense connected followed by a 1 × 1 convolution layer. In addition, to improve the performance in the case of large upsampling factors (e.g., 16×), the depth features are enhanced from coarse to fine. In each frequency-dependent refinement of the depth features, the above iterative subnetwork as well as the residual learning are introduced. The proposed method is tested for the noise-free and noisy cases which compares against 16 state-of-the-art methods. Our experimental results show the improved performances based on the qualitative and quantitative evaluations. Yifan Zuo 0001, Yuming Fang 0001, Ping An 0001, Xiwu Shang, Junnan Yang |
IEEE Trans. Multim. | 3 |
| 2020 | Video intra prediction using convolutional encoder decoder network
Zhipeng Jin, Ping An 0001, Liquan Shen |
Neurocomputing | 2 |
| 2020 | Post-processing for intra coding through perceptual adversarial learning and progressive refinement
Zhipeng Jin, Ping An 0001, Chao Yang 0021, Liquan Shen |
Neurocomputing | 2 |
| 2020 | Screen content image quality assessment based on convolutional neural networks
Xuhao Jiang, Liquan Shen, Linru Zheng, Ping An 0001 |
J. Vis. Commun. Image Represent. | 5 |
| 2020 | Error sensitivity model based on spatial and temporal features
Dezhi Bo, Qiang Wu 0001, Ping An 0001 |
Multim. Tools Appl. | 5 |
| 2020 | Light Field Compression Using Global Multiplane Representation and Two-Step PredictionabstractDue to its spatio-angular structure, light field image allows for a wealth of post-processing techniques like digital refocusing and depth estimation. In order to compress the data of the two domains, the current proposal intends to embed the disparity-based view synthesis method into the decoder. However, predicting each view separately or in local groups means bringing more computational burden to the decoder and destroying the light field structure. Since disparity contains the relationship between all light rays in the light field, the proposed solution is to predict a disparity-based global representation as the first step. In the second step, all the views can be predicted easily based on this representation. In this letter, we use the recently proposed multiplane as the form of this global representation. The experimental results show the effectiveness of the proposed solution, and the better RD performance compared to other schemes especially under low bitrates. Ping An 0001, Xinpeng Huang, Chao Yang 0021, Deyang Liu, Qiang Wu 0001 |
IEEE Signal Process. Lett. | 2 |
| 2020 | Full Reference Light Field Image Quality Evaluation Based on Angular-Spatial CharacteristicabstractThe quality evaluation is an indispensable link in light field (LF) image processing. Most of existing LF objective evaluation indexes do not make effective use of the angular characteristic of LF, so the evaluation results are unsatisfactory. In this letter, the quality evaluation of LF image is constructed based on human visual system (HVS) and LF angular-spatial characteristics. Based on the fact that HVS has different sensitivity to different parallaxes, we assume that LF image quality perceived by human eyes has the optimal parallax range. A dual-fan filter is used to constrain the parallax range. Then, the overall quality of the LF is represented by combining the spatial and angular quality, which performed from the central sub-aperture image and the focus stack, respectively. In addition, because the difference of Gaussian (DoG) operator can simulate the process of extracting texture structure by human eyes. The structural similarity of DoG texture feature is utilized in the spatial quality evaluation. Extensive comparison experiments show that the proposed method is more consistent with the characteristics of LF. Chunli Meng, Ping An 0001, Xinpeng Huang, Chao Yang 0021, Deyang Liu |
IEEE Signal Process. Lett. | 2 |
| 2020 | Low-Complexity CTU Partition Structure Decision and Fast Intra Mode Decision for Versatile Video CodingabstractQuadtree with nested multi-type tree (QTMT) partition structure is an efficient improvement in versatile video coding (VVC) over the quadtree (QT) structure in the advanced high-efficiency video coding (HEVC) standard. With the exception of the recursive QT partition structure, recursive multi-type tree partition is applied to each leaf node, which generates more flexible block sizes. Besides, intra prediction modes are extended from 35 to 67 so as to satisfy various texture patterns. These newly developed techniques achieve high coding efficiency but also result in very high computational complexity. To tackle this problem, we propose a fast intra-coding algorithm consisting of low-complexity coding tree units (CTU) structure decision and fast intra mode decision in this paper. The contributions of the proposed algorithm lie in the following aspects: 1) the new block size and coding mode distribution features are first explored for a reasonable fast coding scheme; 2) a novel fast QTMT partition decision framework is developed, which can determine the partition decision on both QT and multi-type tree with a novel cascade decision structure; and 3) fast intra mode decision with gradient descent search is introduced, while the best initial search point and search step are also investigated in this paper. The simulation results show that the complexity reduction of the proposed algorithm is up to 70% compared to VVC reference software (VTM), and averagely 63% encoding time saving is achieved with 1.93% BDBR increasing. Such results demonstrate that our method yields a superior performance in terms of computational complexity and compression quality compared to the state-of-the-art methods. Hao Yang 0008, Liquan Shen, Xinchao Dong, Ping An 0001, Gangyi Jiang |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2020 | Multi-Scale Frequency Reconstruction for Guided Depth Map Super-Resolution via Deep Residual NetworkabstractThe depth maps obtained by the consumer-level sensors are always noisy in the low-resolution (LR) domain. Existing methods for the guided depth super-resolution, which are based on the pre-defined local and global models, perform well in general cases (e.g., joint bilateral filter and Markov random field). However, such model-based methods may fail to describe the potential relationship between RGB-D image pairs. To solve this problem, this paper proposes a data-driven approach based on the deep convolutional neural network with global and local residual learning. It progressively upsamples the LR depth map guided by the high-resolution intensity image in multiple scales. A global residual learning is adopted to learn the difference between the ground truth and the coarsely upsampled depth map, and the local residual learning is introduced in each scale-dependent reconstruction sub-network. This scheme can restore the depth structure from coarse to fine via multi-scale frequency synthesis. In addition, batch normalization layers are used to improve the performance of depth map denoising. Our method is evaluated in noise-free and noisy cases. A comprehensive comparison against 17 state-of-the-art methods is carried out. The experimental results show that the proposed method has faster convergence speed as well as improved performances based on the qualitative and quantitative evaluations. Yifan Zuo 0001, Qiang Wu 0001, Yuming Fang 0001, Ping An 0001, Liqin Huang |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2020 | Content-Based Light Field Image Compression Method With Gaussian Process RegressionabstractLight field (LF) imaging enables new possibilities for digital imaging, such as digital refocusing, changing of focus plane, changing of viewpoint, scene-depth estimation, and 3D scene reconstruction, by capturing both spatial and angular information of light rays. However, one main problem in dealing with LF data is its sheer volume. In this context, efficient compression methods are needed for such a particular type of content. In this paper, we propose a content-based LF image-compression method with Gaussian process regression to improve the compression efficiency and accelerate the prediction procedure. First, the LF image is fed to the intra-frame codec of HEVC. In the prediction procedure, the prediction units (PUs) are classified as non-homogenous texture units, homogenous texture units, and visually flat units, based on the content property of the LF image. For each category, we design a corresponding Gaussian process regression (GPR)-based prediction method. Moreover, we propose a classification mechanism to exactly decide to which category the current PU belongs, so as to adjust the trade-off between the computational burden and the LF image coding efficiency. Experimental results demonstrate that the proposed LF image compression method is superior to several other state-of-the-art compression methods in terms of different quality metrics. Furthermore, the proposed method can also achieve a good visual quality of views rendered from decoded LF contents. Deyang Liu, Ping An 0001, Wenfa Zhan, Xinpeng Huang, Ali Abdullah Yahya |
IEEE Trans. Multim. | 2 |
| 2019 | Objective Quality Assessment for Light Field Based on Refocus Characteristic
Chunli Meng, Ping An 0001, Xinpeng Huang, Chao Yang 0021 |
ICIG (3) | 2 |
| 2019 | Context-Aware Natural Integration of Advertisement ObjectabstractProduct placement, also called advertisement embedding, is to place some specific products in an image or a video, which may attract consumers to buy their products. However, adding advertisement objects in images is difficult, because where to add the product and how to fuse the background must be concerned. In this paper, to overcome this issue, we present a novel hierarchical framework with conditional generative adversarial network to add advertisement object in all kinds of scene images. The key point of our framework is leaning the relation between surrounding and products. To generate the products more realistic and to make the detail information such characters and logo more clear, we propose local-global discriminator. Our experiment demonstrates the effectiveness of our proposed method on different scene images. Detail codes and results are available in GitHub. https://github.com/happyyuwei/ProductPlacement. Yanhong Ding, Guowei Teng, Yuwei Yao, Ping An 0001, Kai Li 0016 |
ICIP | 4 |
| 2019 | A Novel No-Reference Quality Assessment Model of Tone-Mapped HDR ImageabstractResearch on tone mapping operators (TMOs) attracts more attention recently, which can transform high dynamic range (HDR) images to low dynamic range (LDR) images for visualizing them on the common displays. In this paper, we propose a novel no-reference image quality assessment (IQA) model to evaluate the perceptual quality of tone-mapped images (TMIs). Specifically, local phase congruency (LPC) is first computed to evaluate the image sharpness and some statistical characteristics are extracted on the edge maps to measure the halo effect. Meanwhile, TMIs are transformed to opponent color (OC) space to gain the global image chromaticity and local image contract in the chromatic field. Finally, a regression module is learnt using support vector regression (SVR) to train the mapping function that maps all the features to subjective quality scores. The model shows admirable performance when tested on ESPL-LIVE HDR image database. Liquan Shen, Mingxing Jiang, Linru Zheng, Ping An 0001 |
ICIP | 5 |
| 2019 | Modified Baseline for Light Field StitchingabstractIn traditional 2D image stitching, the baseline method usually means global homography via Direct Linear Transformation (DLT) on inliers. In this paper, a modified baseline method for light field (LF) stitching is proposed to stitch two LFs. The depth map and the center sub-aperture image (SAI) are used to filter the feature points of the entire LF. The global 4D homography is then calculated by DLT to align all SAIs corresponding to the same angular domain coordinates of two LFs. Finally, the improved Markov Random Field (MRF) energy considering the global LF is used to find the seam of 2D SAIs instead of computational 4D graph cut. Experimental results show that the proposed method can effectively stitch the 4D LFs, and preserve the consistency of the angular and spatial domains of the stitched LF compared with implementing 2D image stitching to the corresponding SAIs. Moreover, the method proposed in this paper can easily extend all advanced 2D image stitching methods to 4D LF, so that the acquired LF can have larger field of view and wider applications. Ping An 0001, Xinpeng Huang, Chunli Meng, Qiang Wu 0001 |
VCIP | 2 |
| 2019 | Jointly learning perceptually heterogeneous features for blind 3D video quality assessment
Shuai Yuan 0008, Yun Zhu 0002, Jian Zhang 0002, Ping An 0001 |
Neurocomputing | 5 |
| 2019 | A content-based rate control algorithm for screen content video coding
Liquan Shen, Hao Yang 0008, Ping An 0001 |
J. Vis. Commun. Image Represent. | 4 |
| 2019 | Quality Enhancement Network via Multi-Reconstruction Recursive Residual Learning for Video CodingabstractLossy compression algorithms introduce multiple compression artifacts that severely decrease visual quality. These compression artifacts are highly related to texture contents, and the hierarchical coding units decision structure also brings multi-scale similarity to these artifacts. Current loop filters fail to utilize these characteristics to comprehensively remove compression artifacts. To this end, this letter proposes a novel quality enhancement method by adopting a multi-reconstruction recurrent residual network (MRRN). In particular, a modified recursive residual structure is designed to capture the multi-scale similarity of compression artifact. To effectively enhance frames with uneven noise, a multi-reconstruction structure is proposed, which outputs images with different denoise ratios and adaptively fuses them. Experimental results show that the proposed MRRN can improve coding efficiency up to 15.1% compared with the original loop filter in high-efficiency video coding. Averagely, 6.7%, 7.8%, 7.6% BD-rate reduction is achieved for all intra, low-delay P, and low-delay B, respectively. Meanwhile, as a quality enhancement method performed at encoder side, MRRN also achieves a good balance between coding performance and computational complexity compared to the state-of-the-art methods. Liangwei Yu, Liquan Shen, Hao Yang 0008, Ping An 0001 |
IEEE Signal Process. Lett. | 5 |
| 2019 | SHVC CU Processing Aided by a Feedforward Neural NetworkabstractThe development of multimedia and hardware technologies has led to a great number of industrial video applications, such as virtual reality, high-definition video surveillance, and remote monitoring. As complex communication environments and heterogeneous networks are common in industrial applications, industrial videos are required to support a diverse range of display resolutions and transmission channel capacities. Scalable high-efficiency video coding (SHVC) standards provide the tools to meet this requirement. However, it is highly computationally expensive. Coding complexity has a great impact on SHVC performance in industrial applications. Many of these applications are sensitive to time delay and have limited power. Thus, improvements are required to ensure the practical usability of SHVC encoders. In SHVC encoders, intra/interprediction of variable coding unit (CU) sizes is independently performed for the base and enhancement layers (ELs). There are many interlayer similarities that can be exploited to speed up the procedure for EL coding. In this paper, we propose a feedforward neural network aided model for CU size and mode decisions for SHVC, which utilizes base layer coding information and the coding data of spatiotemporal neighboring CUs to decide which CU sizes or prediction modes can be bypassed for certain EL CUs. Two feedforward neural network based learning models are built for CU classification, which are introduced in the procedures for CU size and mode decisions, respectively. According to the analysis from a large number of video sequences, the representative features are directly extracted from the coding information of previously coded neighboring CUs to avoid computational overheads. After the training is finished, these two models are designed and integrated to build classifiers. Then, two online classification approaches are designed for the CU size and mode decision procedures to classify each CU's type. Finally, different candidate CU sizes and prediction modes are adaptively assigned for each type of CU. This approach outperforms the state-of-the-art fast SHVC/high-efficiency video coding (HEVC) algorithms with approximately 19-42% coding time savings or better compression efficiency, which will be beneficial for the realization of real-time scalable video coding. Liquan Shen, Guorui Feng, Ping An 0001 |
IEEE Trans. Ind. Informatics | 3 |
| 2019 | A Computational Model for Stereoscopic Visual Saliency PredictionabstractDepth information plays an important role in human vision as it provides additional cues that distinguish objects from their backgrounds. This paper explores depth information for analyzing stereoscopic saliency and presents a computational model that predicts stereoscopic visual saliency based on three aspects of human vision: 1) the pop-out effect; 2) comfort zones; and 3) background effects. Through an analysis of these three phenomena, we find that most of the stereoscopic saliency region can be explained. Our model comprises three modules, each describing one aspect of saliency distribution, and a control function that can be used to adjust the three models independently. The relationship between the three models is not mutually exclusive. One, two, or three phenomena may appear in one image. Therefore, to accurately determine which phenomena the image conforms to, we have devised a selection strategy that chooses the appropriate combination of models based on the content of the image. Our approach is implemented within a framework based on the multifeature analysis. The framework considers surrounding regions, color/depth contrast, and points of interest. The selection strategy can improve the performance of the framework. A series of experiments on two recent eye-tracking datasets shows that our proposed method outperforms several state-of-the-art saliency models. Hao Cheng 0003, Jian Zhang 0002, Qiang Wu 0001, Ping An 0001 |
IEEE Trans. Multim. | 4 |
| 2019 | A Modified Just Noticeable Depth Difference Model Built in Perceived Depth SpaceabstractThis paper proposes a modified just noticeable depth difference (JNDD) (JNDiD) model in perceived depth space. The JNDiD model improves the accuracy of current JNDD models by also taking into account the blurriness caused by the change in accommodation when a 3-D object is perceived far from the screen. This change in accommodation is a result of the convergence-accommodation conflict, in which convergence plays a leading role. In the JNDiD model described in this paper, the JNDD threshold is the addition between a base threshold and an additional threshold. Adapted to different blurring effects in three depth regions, the additional threshold is defined as a three-piecewise linear function. The proposed model also attempts to separate the JNDD modeling from the display modeling and viewing conditions in perceived depth space, making it applicable for different types of displays given their specific model parameters. With the help of model parameter transfer functions, the JNDiD model in perceived depth space and the corresponding one in any specific stimulated depth space can be easily transformed to each other. Experimental results demonstrate the effectiveness and superiority of the JNDiD model in perceived depth space. Ping An 0001, Liquan Shen, Kai Li 0016 |
IEEE Trans. Multim. | 2 |
| 2019 | No-Reference Quality Assessment for Screen Content Images Based on Hybrid Region Features FusionabstractResearch on screen content images (SCIs) attracts more attention as they are highly applied to image- and video-centric applications on mobile and other devices. It is important to develop an efficient image-quality assessment (IQA) method for SCIs because IQA can guide and optimize various image-processing methods for SCIs and improve user experience. In this paper, we propose a no-reference objective assessment model for SCIs including SCIs segmentation and the analysis of local and global perceptual feature representations. Since the human visual system is highly sensitive to sharp edges that are commonly encountered in SCIs, we utilize the variance of local standard deviation, which is a noise robust index to distinguish the sharp edge patches (SEPes) and non-SEPes of SCIs. For SEPes, we perform two kinds of feature extractions. First, the entropy and contrast features are extracted with a gray-level co-occurrence matrix, which are highly perceptive of microstructural change. Second, the local phase coherence is utilized to capture the loss in sharpness. Then, average pooling is adopted to fuse features obtained from all of the SEPes to represent the local features. We further combine local features with global features that are derived using the BRISQUE method as the hybrid region (HR)-based features. Finally, a regression module is learned using support vector regression to train the mapping function that maps HR-based features to subjective quality scores. Experimental results on the screen image-quality assessment database show that the proposed method can achieve better performance in visual-quality prediction for SCIs than the performance achieved by state-of-the-art methods. Linru Zheng, Liquan Shen, Jianan Chen 0001, Ping An 0001, Jun Luo 0006 |
IEEE Trans. Multim. | 4 |
| 2019 | Low-Complexity Scalable Extension of the High-Efficiency Video Coding (SHVC) Encoding SystemabstractThe scalable extension of the high-efficiency video coding (SHVC) system adopts a hierarchical quadtree-based coding unit (CU) that is suitable for various texture and motion properties of videos. Currently, the test model of SHVC identifies the optimal CU size by performing an exhaustive quadtree depth-level search, which achieves a high compression efficiency at a heavy cost in terms of the computational complexity. However, many interactive multimedia applications, such as remote monitoring and video surveillance, which are sensitive to time delays, have insufficient computational power for coding high-definition (HD) and ultra-high-definition (UHD) videos. Therefore, it is important, yet challenging, to optimize the SHVC coding procedure and accelerate video coding. In this article, we propose a fast CU quadtree depth-level decision algorithm for inter-frames on enhancement layers that is based on an analysis of inter-layer, spatial, and temporal correlations. When motion/texture properties of coding regions can be identified early, a fast algorithm can be designed for adapting CU depth-level decision procedures to video contents and avoiding unnecessary computations during CU depth-level traversal. The proposed algorithm determines the motion activity level at the treeblock size of the hierarchical quadtree by utilizing motion vectors from its corresponding blocks at the base layer. Based on the motion activity level, neighboring encoded CUs that have larger correlations are preferentially selected to predict the optimal depth level of the current treeblock. Finally, two parameters, namely, the motion activity level and the predicted CU depth level, are used to identify a subset of candidate CU depth levels and adaptively optimize CU depth-level decision processes. The experimental results demonstrate that the proposed scheme can run approximately three times faster than the most recent SHVC reference software, with a negligible loss of compression efficiency. The proposed scheme is efficient for all types of scalable video sequences under various coding conditions and outperforms state-of-the-art fast SHVC and HEVC algorithms. Our scheme is a suitable candidate for interactive HD/UHD video applications that are expected to operate in real-time and power-constrained scenarios. Liquan Shen, Ping An 0001, Guorui Feng |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2018 | Quality Enhancement for Intra Frame Coding Via Cnns: An Adversarial ApproachabstractLossy compression is an indispensable technique in image/video processing, due to its highly desirable ability of reducing the huge data volume. However, lossy compression introduces complex compression artifacts. To reduce these artifacts, post-processing techniques have been extensively studied. In this paper, we propose a novel post-processing technique using multi-level progressive refinement network via an adversarial training approach, called MPRGAN, for artifacts reduction and coding efficiency improvement in intra frame coding. Furthermore, our network generates multi-level residues in one feed-forward pass through the progressive reconstruction. This coarse-to-fine work fashion, which makes our network have high flexibility, can make trade-off between enhanced quality and computational complexity. Thereby facilitates the resource-aware applications. Extensive evaluations on benchmark datasets verify the superiority of our proposed MPRGAN model over the latest state-of-the-art methods with fast deployment running speed. Zhipeng Jin, Ping An 0001, Chao Yang 0021, Liquan Shen |
ICASSP | 2 |
| 2018 | View Synthesis for Light Field Coding Using Depth EstimationabstractLight Field (LF) image captured by plenoptic camera can record richer scenario information from our world. But a huge number of Sub-Aperture Images (SAIs) from LF image results in great challenges for coding LFI. Therefore, we choose a subset of SAIs as multi-view video, and encode them with their corresponding depth maps using Multi-view Video plus Depth (MVD) structure. Due to lack of depth maps for SAIs, we propose a cost function to determine the depth map for each SAI preliminarily based on horizontal and vertical Epipolar Plane Images (EPI), respectively. Then an SAI-guided depth enhancement algorithm is designed to optimize the estimated depth maps. Since those unselected SAIs have not been encoded but have been synthesized using the specific texture image and depth map, our LF image coding method can naturally achieve bitrates reduction dramatically with a good performance and outperform other algorithms significantly. Xinpeng Huang, Ping An 0001, Liang Shan 0004, Liquan Shen |
ICME | 2 |
| 2018 | No-reference stereo image quality assessment by learning gradient dictionary-based color visual characteristicsabstractIn this paper, we propose a no-reference (NR) stereo image quality assessment metric by learning gradient dictionary-based color visual characteristics. To be specific, firstly, since human eyes are highly sensitive to the structure of images, the gradient magnitude (GM) and gradient orientation (GO) are extracted from left and right views of stereo image, meanwhile, the difference map is obtained. Considering the influence of color distortion, images are decomposed into RGB channels to be processed respectively, and we get the local gradient of the color image by adding up the RGB gradient vectors. Constructively, the gradient dictionary is generated, which is different from traditional image dictionary. All quality-aware features are extracted by joint sparse representation. Afterwards, to avoid over-fitting, the principal component analysis (PCA) is applied to optimize the quality-aware features. Finally, all features are fed into the trained support vector regression (SVR) model to predict the objective score. The experimental results show that the proposed metric always achieves high consistency with human subjective assessment for both symmetric and asymmetric distortions. Jialu Yang, Ping An 0001, Jian Ma 0005, Kai Li 0016, Liquan Shen |
ISCAS | 2 |
| 2018 | Light Field Image Sparse Coding via CNN-Based EPI Super-ResolutionabstractThis paper proposes a novel light field (LF) image compression scheme by super resolving the epipolar plane image (EPI) via convolutional neural network (CNN). In the scheme, we first decompose the LF image into sub-aperture images (SAIs), and only one quarter of them are compressed on the encoding side to reduce the bitrate. On the decoding side, we use these selected SAIs to reconstruct the entire LF by taking advantage of the special structure of EPI. The low-resolution EPIs generated from the sparse SAIs are super resolved by using deep residual network and the output high-resolution EPIs are used to rebuild the dense SAIs. Experimental results show the superior performance of our scheme, which achieve 1.46 dB quality improvement and 35.85 percent bit rate reduction on average compared with the typical pseudo-sequence-based coding method. Jinbo Zhao, Ping An 0001, Xinpeng Huang, Liang Shan 0004 |
VCIP | 2 |
| 2018 | Hybrid linear weighted prediction and intra block copy based light field image coding
Deyang Liu, Ping An 0001, Liquan Shen |
Multim. Tools Appl. | 2 |
| 2018 | Scalable coding of 3D holoscopic image by using a sparse interlaced view image set and disparity map
Deyang Liu, Ping An 0001, Chao Yang 0021, Liquan Shen, Kai Li 0016 |
Multim. Tools Appl. | 2 |
| 2018 | Joint binocular energy-contrast perception for quality assessment of stereoscopic images
Jian Ma 0005, Ping An 0001, Liquan Shen, Kai Li 0016 |
Signal Process. Image Commun. | 2 |
| 2018 | Efficient screen content intra coding based on statistical learning
Hao Yang 0008, Liquan Shen, Ping An 0001 |
Signal Process. Image Commun. | 3 |
| 2018 | Bivariate analysis of 3D structure for stereoscopic image quality assessment
Liquan Shen, Ping An 0001 |
Signal Process. Image Commun. | 3 |
| 2018 | Fast Intra Coding of High Dynamic Range Videos in SHVCabstractCompared with the conventional standard dynamic range (SDR) content, high dynamic range (HDR) content supplies viewers with more immersive experience by offering a much higher range of luminance. Most of current consumer devices cannot afford to this emerging technology, and content providers decide to create both an HDR version and an SDR version of the same video. In this letter, scalable high efficiency video coding (HEVC) scalable extension of HEVC (SHVC) serves as the coding framework where the base layer (BL) is an 8-b SDR version and the enhancement layer (EL) is a 12-b HDR version. Recently, many fast coding algorithms for SDR videos are proposed, and there is an urgent demand for fast coding algorithms for EL HDR videos. With the coding information of the BL SDR videos, this letter proposes a fast algorithm to reduce the complexity of intra coding for EL HDR videos. First, depth information of neighboring coding tree units (CTUs) in the HDR version and the colocated CTU in the SDR version is used for early coding unit (CU) depth determination. Moreover, four classifiers are trained to predict the CTU depth range. Two classifiers are trained for CTUs in frames with a high average luma, and another two classifiers are used for CTUs in frames with a low average luma. Experimental results show that the proposed algorithm achieves 43% encoding time saving on average, with only a 0.54% Bjøntegaard delta bit rate (BDBR) increase compared to the original SHVC test model. Guoliang Fu, Liquan Shen, Hao Yang 0008, Xiangyu Hu 0003, Ping An 0001 |
IEEE Signal Process. Lett. | 5 |
| 2018 | Explicit Edge Inconsistency Evaluation Model for Color-Guided Depth Map EnhancementabstractColor-guided depth enhancement is used to refine depth maps according to the assumption that the depth edges and the color edges at the corresponding locations are consistent. In methods on such low-level vision tasks, the Markov random field (MRF), including its variants, is one of the major approaches that have dominated this area for several years. However, the assumption above is not always true. To tackle the problem, the state-of-the-art solutions are to adjust the weighting coefficient inside the smoothness term of the MRF model. These methods lack an explicit evaluation model to quantitatively measure the inconsistency between the depth edge map and the color edge map, so they cannot adaptively control the efforts of the guidance from the color image for depth enhancement, leading to various defects such as texture-copy artifacts and blurring depth edges. In this paper, we propose a quantitative measurement on such inconsistency and explicitly embed it into the smoothness term. The proposed method demonstrates promising experimental results compared with the benchmark and state-of-the-art methods on the Middlebury ToF-Mark, and NYU data sets. Yifan Zuo 0001, Qiang Wu 0001, Jian Zhang 0002, Ping An 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2018 | Efficient Intra Mode Selection for Depth-Map Coding Utilizing Spatiotemporal, Inter-Component and Inter-View Correlations in 3D-HEVCabstract3D-high efficiency video coding (HEVC) is developed for the compression of the multi-view video plus depth format, which is based on the latest generation of video coding standard, HEVC. It further adopts several new intra prediction modes, depth-modeling modes (DMMs) in intra candidate modes for a better representation of edges in depth maps, which introduces a drastic increase in the computational complexity. The procedure of depth intra mode decision together with DMMs and existing intra modes is a very time consuming part due to huge complexity of full rate distortion (RD) cost calculation. In this paper, a low complexity intra mode selection algorithm is proposed to reduce complexity of depth intra prediction in both intra-frames and inter-frames. An experimental analysis is first performed to study the inter-view correlation and the inter-component (texture video and its associated depth) correlation in intra coding information such as the intra mode and RD cost. All intra modes available in 3D-HEVC are classified into three activity classes assigned with different mode-weight factors, and the coding mode complexity of a coding unit (CU) is defined according to the intra mode information from available spatiotemporal, inter-view, and inter-component neighboring coded CUs. The coding mode complexity analysis is utilized to assign different candidate intra modes for different types of CUs. The optimal intra prediction mode and the RD cost value in current CU depth level are further used to skip unnecessary intra prediction sizes. Experimental results show that the proposed fast depth intra coding algorithm achieves 61% complexity reduction on intra prediction, while incurring a 0.2% Bjontegaard metric increase for coded and synthesized views compared to the test model of 3D-HEVC. Liquan Shen, Kai Li 0016, Guorui Feng, Ping An 0001, Zhi Liu 0003 |
IEEE Trans. Image Process. | 4 |
| 2018 | Minimum Spanning Forest With Embedded Edge Inconsistency Measurement Model for Guided Depth Map EnhancementabstractGuided depth map enhancement based on Markov Random Field (MRF) normally assumes edge consistency between the color image and the corresponding depth map. Under this assumption, the low-quality depth edges can be refined according to the guidance from the high-quality color image. However, such consistency is not always true, which leads to texture-copying artifacts and blurring depth edges. In addition, the previous MRF-based models always calculate the guidance affinities in the regularization term via a non-structural scheme which ignores the local structure on the depth map. In this paper, a novel MRF-based method is proposed. It computes these affinities via the distance between pixels in a space consisting of the Minimum Spanning Trees (Forest) to better preserve depth edges. Furthermore, inside each Minimum Spanning Tree, the weights of edges are computed based on explicit edge inconsistency measurement model, which significantly mitigates texture-copying artifacts. To further tolerate the effects caused by noise and better preserve depth edges, a bandwidth adaption scheme is proposed. Our method is evaluated for depth map super-resolution and depth map completion problems on synthetic and real datasets including Middlebury, ToF-Mark and NYU. A comprehensive comparison against 16 state-of-the-art methods is carried out. Both qualitative and quantitative evaluation present the improved performances. Yifan Zuo 0001, Qiang Wu 0001, Jian Zhang 0002, Ping An 0001 |
IEEE Trans. Image Process. | 4 |
| 2017 | Coding of 3D holoscopic image by using spatial correlation of rendered view imagesabstractHoloscopic imaging is a prospective acquisition and display solution for providing natural and fatigue-free 3D visualization. However, large amount of data is required to represent the 3D holoscopic content. Therefore, efficient coding schemes for this particular type of image are needed. In this paper, an effective coding scheme is proposed by exploring the spatial correlation among the view images with different perspectives rendered from 3D holoscopic image. We utilize the interlaced view image to descript such spatial correlation. A linear prediction method is used on the interlaced view image instead of the original holoscopic image directly. Experimental results show that the proposed coding scheme performs better than HEVC intra standard and screen content coding extension of HEVC with around 2.41dB and 0.42 dB average quality improvement respectively. Deyang Liu, Ping An 0001, Chao Yang 0021, Liquan Shen |
ICASSP | 2 |
| 2017 | Sparse Time-Varying Graphs for Slide Transition Detection in Lecture Videos
Zhijin Liu, Kai Li 0016, Liquan Shen, Ping An 0001 |
ICIG (1) | 4 |
| 2017 | An efficient intra coding algorithm based on statistical learning for screen content codingabstractScreen content has different characteristics compared with natural content captured by cameras. To achieve more efficient compression, some new coding tools have been developed in the High Efficiency Video Coding (HEVC) Screen Content Coding (SCC) Extension, which also increase the computational complexity of encoder. In this paper, complexity analysis are first conducted to explore the distribution of complexities. Then, two classification trees, including early coding units (CU) partition tree (EPT) and CU content classification tree (CCT), are designed based on statistical characteristics and coding information. EPT is used to decide whether the CU skip the mode decision process of current depth level and CCT is used to classify the blocks into either natural blocks or screen blocks. Natural blocks will skip screen coding modes and screen blocks skip normal intra modes. Experimental results show the proposed algorithm can save 49% encoding time with 2.7% BD-rate increase on average for All Intra configuration under the SCC common test condition. Hao Yang 0008, Liquan Shen, Ping An 0001 |
ICIP | 3 |
| 2017 | Minimum spanning forest with embedded edge inconsistency measurement for color-guided depth map upsamplingabstractColor-guided depth map up-sampling, such as Markov-Random-Field-based (MRF-based) methods, is a popular depth map enhancement solution, which normally assumes edge consistency between color image and corresponding depth map. It calculates the coefficients of smoothness term in MRF according to such assumption. However, such consistency is not always true which leads to texture-copying artifacts and blurring depth edges. In this paper, we propose a novel coefficient computing scheme for smoothness term in MRF which is based on the distance between pixels in the Minimum Spanning Trees (Forest) to better preserve depth edges. The explicit edge inconsistency measurement is embedded into weights of edges in Minimum Spanning Trees, which significantly mitigates texture-copying artifacts. The proposed method is evaluated on Middlebury datasets and ToF-Mark datasets which demonstrates improved results compared with state-of-the-art methods. Yifan Zuo 0001, Qiang Wu 0001, Jian Zhang 0002, Ping An 0001 |
ICME | 4 |
| 2017 | CNN oriented fast QTBT partition algorithm for JVET intra codingabstractIn this paper, a novel fast coding unit depth decision algorithm based on convolution neural network is presented for JVET future video coding. JVET employs quad-tree plus binary-tree (QTBT) block partitioning structure, which can support much more flexibility for coding units partition shapes, and improve the coding performance significantly than the HEVC standard. However, the flexible partitioning structure also introduces a tremendous computation complexity. To address this issue, we model the QTBT partition depth range as a multi-class classification problem, and try to predict the depth range of 32×32 block directly, rather than to judge split or not at each depth level. To the best of our knowledge, it is the first framework to formulate the QTBT partition range as a multi classification task, and optimized by an end-to-end learning model. For training optimization, we design an objective function consists of class penalty term and L2 HingeLoss function, which leverage the characteristics of category settings, can further boost the classification accuracy. Experimental results demonstrate the effectiveness of our proposed method, which can achieve 42.80% complexity reduction with only 0.65% Bjontegaard Delta bitrate (BD-rate) increase. Zhipeng Jin, Ping An 0001, Liquan Shen, Chao Yang 0021 |
VCIP | 2 |
| 2017 | SSIM-based binocular perceptual model for quality assessment of stereoscopic imagesabstractIn this paper, we propose a novel full reference stereoscopic image quality assessment (FR-SIQA) metric by utilizing SSIM-based binocular perceptual model. The goal is to predict the perceptual quality of a stereoscopic image via jointly considering the qualities of cyclopean image and the difference image. Specifically, we first apply the contrast sensitivity filtering to both the reference and distorted stereo pairs. Constructively, a new cyclopean image is generated by considering binocular perceptual model and binocular rivalry simultaneously. Finally, the overall quality score of a testing stereoscopic image is predicted by combining the qualities of its cyclopean image and difference image. Experimental results show that the proposed metric achieves high consistency with human subjective assessment and outperforms several the state-of-the-art FR-SIQA methods. Jian Ma 0005, Ping An 0001, Liquan Shen, Kai Li 0016, Jialu Yang |
VCIP | 2 |
| 2017 | Bivariate statistics and binocular energy induced stereo-pair quality evaluatorabstractWith the flourishment of 3D content, stereoscopic image quality assessment (SIQA) becomes an urgent issue in image processing field. In this paper, a new blind SIQA method is proposed based on bivariate natural scene statistics (NSS) model that is conducted on the binocular Gabor energy response. Specifically, Gabor responses of two views are combined using a binocular energy model. Then bivariate statistics of the jointly spatially adjacent responses of the fused Gabor energy are calculated for feature extraction. Experimental results on LIVE 3D Image Quality Database demonstrate the promising performance of the proposed method. Liquan Shen, Ping An 0001 |
VCIP | 3 |
| 2017 | A stereoscopic image quality assessment model based on independent component analysis and binocular fusion property
Xianqiu Geng, Liquan Shen, Kai Li 0016, Ping An 0001 |
Signal Process. Image Commun. | 4 |
| 2017 | Bit allocation for 3D video coding based on lagrangian multiplier adjustment
Chao Yang 0021, Ping An 0001, Deyang Liu, Liquan Shen, Kai Li 0016 |
Signal Process. Image Commun. | 2 |
| 2016 | Depth map coding based on virtual view qualityabstractMulti-view video plus depth (MVD) is a 3D video representation. In MVD, the depth map provides the scene distance information and is used to render the virtual view through Depth Image Based Rendering (DIBR) technique. The depth map coding error will induce distortion in the rendered virtual views. This paper proposes a mathematic model that can estimate the synthesized virtual view distortion induced by depth map compression, and the model is employed to the rate distortion optimization (RDO) in the depth map coding. Based on the rendered virtual view quality, a Lagrangian optimization adjustment scheme at Coding Unit (CU) level is proposed to improve the depth map encoding efficiency. Experimental results demonstrate that the proposed method can improve the BD-PSNR of virtual view for 0.62 dB, and the encoding complexity reduces compared with the view synthesis optimization (VSO) technique in the 3D-HEVC Test Model (HTM). Chao Yang 0021, Ping An 0001, Deyang Liu, Liquan Shen |
ICASSP | 2 |
| 2016 | Explicit measurement on depth-color inconsistency for depth completionabstractColor-guided depth completion is to refine depth map through structure light sensing by filling missing depth structure and de-nosing. It is based on the assumption that depth discontinuity and color edge at the corresponding location are consistent. Among all proposed methods, MRF-based method including its variants is one of major approaches. However, the assumption above is not always true, which causes texture-copy and depth discontinuity blurring artifacts. The state-of-the-art solutions usually are to modify the weighting inside smoothness term of MRF model. Because there is no any method explicitly considering the inconsistency occurring between depth discontinuity and the corresponding color edge, they cannot adaptively control the effect of guidance from color image when completing depth map. In this paper, we propose quantitative measurement on such inconsistency and explicitly embed it into weighting value of smoothness term. The proposed method is evaluated on NYU Kinect datasets and demonstrates promising results. Yifan Zuo 0001, Qiang Wu 0001, Ping An 0001, Jian Zhang 0002 |
ICIP | 3 |
| 2016 | Explicit modeling on depth-color inconsistency for color-guided depth up-samplingabstractColor-guided depth up-sampling is to enhance the resolution of depth map according to the assumption that the depth discontinuity and color image edge at the corresponding location are consistent. Through all methods reported, MRF including its variants is one of major approaches, which has dominated in this area for several years. However, the assumption above is not always true. Solution usually is to adjust the weighting inside smoothness term in MRF model. But there is no any method explicitly considering the inconsistency occurring between depth discontinuity and the corresponding color edge. In this paper, we propose quantitative measurement on such inconsistency and explicitly embed it into weighting value of smoothness term. Such solution has not been reported in the literature. The improved depth up-sampling based on the proposed method is evaluated on Middlebury datasets and ToFMark datasets and demonstrate promising results. Yifan Zuo 0001, Qiang Wu 0001, Jian Zhang 0002, Ping An 0001 |
ICME | 4 |
| 2016 | Using independent component analysis and binocular combination for stereoscopic image quality assessmentabstractIn this paper, a full reference stereoscopic image quality assessment (FR-SIQA) method is proposed based on independent component analysis (ICA) and binocular combination. Image features that reflect the responds of simple cells in the cortex are extracted by ICA-based algorithm. Both image feature similarity (IFS) and local luminance consistency (LLC) are calculated to measure the structure and brightness distortions, respectively. To simulate the binocular fusion properties, the energy of image features and the global relative luminance information are selected as the basic of binocular combination to fuse the right-left IFS and LLC into a final index. Experimental results demonstrate that the proposed algorithm achieves high consistency with subjective assessment on two public available 3D image quality assessment databases. Xianqiu Geng, Liquan Shen, Ping An 0001, Zhi Liu 0003 |
VCIP | 3 |
| 2016 | Just noticeable disparity difference model for 3D displaysabstractBased on the related psychological and physiological advancement, a just noticeable disparity difference (JNDiD) model for 3D displays under the assumption that eyes converge at the virtual object is presented in this paper. Specifically, according to whether the surface of the virtual object is perceived to be blurred, the perceived depth space is divided into three regions. Furthermore, a three-phase linear function considering the accommodation convergence mismatch in different depth regions is employed to build the JNDiD model. Experimental results demonstrate the effectiveness and superiority of our method compared with the state-of-the-art just noticeable depth difference (JNDD) model for 3D displays. Ping An 0001, Liquan Shen, Kai Li 0016, Nina Feng |
VCIP | 2 |
| 2016 | Method to quality assessment of stereo imagesabstractIn this paper, we propose a novel full reference image quality assessment model for stereo images. To improve the characteristics of human vision system (HVS) for stereo quality analysis, this model exploits mechanism of the HVS from three aspects: 1) apply contrast sensitivity function (CSF) filtering on the two monocular views. 2) the conventional 2D MAD (most apparent distortion) algorithm is applied on the two monocular views, and then the combined binocular quality is estimated via a weighted sum of the estimates from two stages. In the first stage, the weights are determined based on binocular phase congruency measure. In the second stage, the weights are determined based on a block-based contrast measure. 3) combine the quality from the two stages into a single estimate of overall perceived quality of stereo images. Experimental results on three public benchmark databases show that the proposed model achieves significantly higher consistency with subjective scores. Jian Ma 0005, Ping An 0001 |
VCIP | 2 |
| 2016 | Parallax-aware local alignment for image stitching under large occlusion/disocclusionabstractThis paper presents a parallax-aware image stitching approach under large occlusion/disocclusion. Different from previous research, we explore the image stitching issue in a parallax-aware perspective via local alignment. Specifically, we first label each feature point with a probability of being large parallax by developing a graph-based optimization framework. Afterwards, an integer programming model is built to pick out a group of feature matches free from parallax in a local region. Finally, by enforcing a stitching seam passing through such a locally aligned area, we are able to generate a high-quality stitching result under large parallax. Experimental results demonstrate the effectiveness and superiority of the proposed parallax-aware approach. Yangxin Wang, Kai Li 0016, Ping An 0001, Liquan Shen, Xuemei Zou |
VCIP | 3 |
| 2016 | 3D holoscopic image coding scheme using HEVC with Gaussian process regression
Deyang Liu, Ping An 0001, Chao Yang 0021, Liquan Shen |
Signal Process. Image Commun. | 2 |
| 2016 | Fast depth map coding based on virtual view quality
Chao Yang 0021, Ping An 0001, Liquan Shen, Nina Feng |
Signal Process. Image Commun. | 2 |
| 2015 | Depth upsampling method via Markov random fields without edge-misaligned artifactsabstractRecently, the widely use of time-of-flight sensors captures depth information for dynamic scenes in real time, which promotes the developing of many 3D image or video processing applications. However, such depth maps are noisy and have low resolutions. In this paper, we propose an edge-based depth map super-resolution method via solving a labeling optimization problem in MRF. The inputs are low quality depth map and the according high-resolution color image. The proposed method not only avoids the texture-copy artifacts, but also preserves the edges of depth which do not exist in the color image. We compare our algorithm with the state of the art on the benchmark dataset. The experimental results prove the validity and robustness of our approach. Yifan Zuo 0001, Ping An 0001, Zhaoyang Zhang 0002 |
ICIP | 2 |
| 2015 | Virtual view distortion estimation for depth map codingabstractMulti-view video plus depth (MVD) format is a three-dimensional (3D) video representation. The depth map in MVD provides the scene geometry information and is used to render the virtual view through Depth Image Based Rendering (DIBR). In this paper, a virtual view distortion estimation function based on the characteristics of both texture image and depth map is proposed which can estimate virtual view distortion induced by depth map compression accurately, and the function is implemented to the Rate Distortion Optimization (RDO) in the depth map coding. Compared with the View Synthesis Optimization (VSO) in 3D-HEVC Test Model (HTM) reference software, the experimental results demonstrate that the proposed method can improve the BD-PSNR of virtual view for 0.26 dB on average, and the encoding time has reduced for 31% on average due to the low complexity of the proposed function. Chao Yang 0021, Ping An 0001, Deyang Liu, Liquan Shen |
VCIP | 2 |
| 2015 | Fast TU size decision algorithm for HEVC encoders using Bayesian theorem detection
Liquan Shen, Zhaoyang Zhang 0002, Xinpeng Zhang 0001, Ping An 0001, Zhi Liu 0003 |
Signal Process. Image Commun. | 4 |
| 2015 | A 3D-HEVC Fast Mode Decision Algorithm for Real-Time Applicationsabstract3D High Efficiency Video Coding (3D-HEVC) is an extension of the HEVC standard for coding of multiview videos and depth maps. It inherits the same quadtree coding structure as HEVC for both components, which allows recursively splitting into four equal-sized coding units (CU). One of 11 different prediction modes is chosen to code a CU in inter-frames. Similar to the joint model of H.264/AVC, the mode decision process in HM (reference software of HEVC) is performed using all the possible depth levels and prediction modes to find the one with the least rate distortion cost using a Lagrange multiplier. Furthermore, both motion estimation and disparity estimation need to be performed in the encoding process of 3D-HEVC. Those tools achieve high coding efficiency, but lead to a significant computational complexity. In this article, we propose a fast mode decision algorithm for 3D-HEVC. Since multiview videos and their associated depth maps represent the same scene, at the same time instant, their prediction modes are closely linked. Furthermore, the prediction information of a CU at the depth level X is strongly related to that of its parent CU at the depth level X-1 in the quadtree coding structure of HEVC since two corresponding CUs from two neighboring depth levels share similar video characteristics. The proposed algorithm jointly exploits the inter-view coding mode correlation, the inter-component (texture-depth) correlation and the inter-level correlation in the quadtree structure of 3D-HEVC. Experimental results show that our algorithm saves 66% encoder runtime on average with only a 0.2% BD-Rate increase on coded views and 1.3% BD-Rate increase on synthesized views. Liquan Shen, Ping An 0001, Zhaoyang Zhang 0002, Qianqian Hu, Zhengchuan Chen |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2012 | Fast Segment-Based Algorithm for Multi-view Depth Map Generation
Yifan Zuo 0001, Ping An 0001, Qiuwen Zhang, Zhaoyang Zhang 0002 |
ICIC (2) | 2 |
| 2011 | An improved depth map estimation for coding and view synthesisabstractInaccuracy depth estimation may influence on depth coding and virtual view rendering in the free-viewpoint television (FTV) system, an improved depth map estimation is proposed to solve the problem for coding and view synthesis. Firstly, check the consistency of initial depth, and the influence of initial miss-matches is minimized by introduction of an additional adaptive matching error selection that penalizes the unreliable matches. Then according to certain criteria, the multi-reference depth maps are merged into one disparity map to improve the quality of disparity map. Finally, a multilateral filtering is used to preserve details in the depth map and simultaneously smooth the depths in occluded areas at object boundary, less texture and discontinuity regions. Experimental results show a significant improvement of the initial input depth maps and coding efficiency, as well as a reduction of view synthesis artifacts. Qiuwen Zhang, Ping An 0001, Liquan Shen, Zhaoyang Zhang 0002 |
ICIP | 2 |
| 2011 | Efficient rendering distortion estimation for depth map compressionabstractA depth map represents three-dimensional (3D) scene information and is used to synthesize virtual views in 3D video. Since the quality of synthesized virtual views highly depends on the quality of depth map, efficient depth compression is crucial to realize the 3D video system. However compressing depth map using existing video coding techniques yields unacceptable distortions while rendering virtual views. To solve this problem, we propose an efficient depth map compression method for the view rendering, a novel distortion metric base on view rendering distortions instead of distortion of depth map itself. First, we derive relationships between distortions in coded depth map and rendered view. Then, a region based video characteristics distortion model is proposed for precisely estimation distortion in view synthesis. Finally, experimental results have shown that 1.8 dB coding gain in terms of PSNR and subjective quality improvement of synthesized views are achieved by the proposed method. Qiuwen Zhang, Ping An 0001, Zhaoyang Zhang 0002 |
ICIP | 2 |
| 2011 | Low-Complexity Mode Decision for MVCabstractThe finalized international standard for multiview video coding (MVC) is an extension of H.264. In the joint model of MVC, variable size motion estimation (ME) and disparity estimation (DE) are introduced to achieve the highest coding efficiency with the cost of very high computational complexity. A low complexity mode decision algorithm is proposed to reduce complexity of ME and DE. An experimental analysis is performed to study inter-view correlation in the coding information such as the prediction mode and rate-distortion (RD) cost. Based on the correlation, we propose four efficient mode decision techniques, including early SKIP mode decision, adaptive early termination, fast mode size decision, and selective intra coding in inter frame. Experimental results show that the proposed algorithm can significantly reduce computational complexity of MVC while maintaining almost the same RD performance. Liquan Shen, Zhi Liu 0003, Ping An 0001, Zhaoyang Zhang 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2010 | An adaptive early termination of mode decision using inter-layer correlation in scalable video codingabstractThe scalable video coding (SVC) standard adopts the variable size motion estimation (ME) to select the best coding mode for each macroblock (MB). Although this technique achieves the highest possible coding efficiency, it results in extremely large computation complexity which obstructs SVC from the practical application. In this paper, we propose an adaptive early termination of fast mode decision algorithm in SVC. It makes use of the coding information of spatial neighbor MBs and the corresponding MBs in base layer to early terminate the mode decision procedure. Experimental results show that the proposed fast mode decision algorithm can achieve computational saving up to 67% with no significant loss of rate distortion (RD) performance. Liquan Shen, Zhi Liu 0003, Ping An 0001, Zhaoyang Zhang 0002 |
ICIP | 3 |
| 2010 | Rate control algorithm based on frame complexity estimation for MVCabstractRate control has not been well studied for multi-view video coding (MVC). In this paper, we propose an efficient rate control algorithm for MVC by improving the quadratic rate-distortion (R-D) model, which reasonably allocate bit-rate among views based on correlation analysis. The proposed algorithm consists of four levels for rate bits control more accurately, of which the frame layer allocates bits according to frame complexity and temporal activity. Extensive experiments show that the proposed algorithm can efficiently implement bit allocation and rate control according to coding parameters. Tao Yan 0003, Ping An 0001, Liquan Shen, Zhaoyang Zhang 0002 |
VCIP | 2 |
| 2010 | Early SKIP mode decision for MVC using inter-view correlation
Liquan Shen, Zhi Liu 0003, Tao Yan 0003, Zhaoyang Zhang 0002, Ping An 0001 |
Signal Process. Image Commun. | 5 |
| 2010 | View-Adaptive Motion Estimation and Disparity Estimation for Low Complexity Multiview Video CodingabstractThe emerging international standard for multiview video coding (MVC) is an extension of H.264/advanced video coding. In the joint mode of MVC, both motion estimation (ME) and disparity estimation (DE) are included in the encoding process. This achieves the highest coding efficiency but requires a very high computational complexity. In this letter, we propose a fast ME and DE algorithm that adaptively utilizes the inter-view correlation. The coding mode complexity and the motion homogeneity of a macroblock (MB) are first analyzed according to the coding modes and motion vectors from the corresponding MBs in the neighbor views, which are located by means of global disparity vector. According to the coding mode complexity and the motion homogeneity, the proposed algorithm adjusts the search strategies for different types of MBs in order to perform a precise search according to video content. Experimental results demonstrate that the proposed algorithm can save 85% computational complexity on average, with negligible loss of coding efficiency. Liquan Shen, Zhi Liu 0003, Tao Yan 0003, Zhaoyang Zhang 0002, Ping An 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2009 | Fast mode decision for multiview video codingabstractIn the draft of multi-view coding (MVC), variable size motion estimation and disparity estimation are employed to select the best coding mode for each macroblock. These techniques achieve the highest possible coding efficiency, but they result in extremely large computation complexity which obstructs MVC from practical application. This paper proposes a fast mode size decision algorithm for MVC in inter-frame coding. It makes use of the mode distribution correlation between neighbor views to deduct the executions of unnecessary modes. Experimental results show that the proposed fast mode decision algorithm reduces the computational complexity significantly with negligible coding efficiency. Liquan Shen, Tao Yan 0003, Zhi Liu 0003, Zhaoyang Zhang 0002, Ping An 0001 |
ICIP | 5 |
| 2004 | Intermediate View Synthesis from Stereoscopic Videoconference Images
Chaohui Lu, Ping An 0001, Zhaoyang Zhang 0002 |
ICCSA (4) | 2 |