Xinpeng Huang

dblp:131/8186 · DBLP profile ↗
← Back
47ranked-venue papers
4as first author
40since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 38 · 3 first-author · 31 since 2021Artificial intelligence and machine learning · 6 · 6 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 VIQ-360: A New Viewport-Based Omnidirectional Image Quality Assessment Database
Xiangxu Yu, Chao Yang 0021, Xinpeng Huang, Ping An 0001
QoMEX4
2026 A Subjective Quality Database for Human-AI Co-Created Images
Xiangxu Yu, Jiyan Tong, Chao Yang 0021, Xinpeng Huang, Ping An 0001
QoMEX6
2026 DMGNet: Discriminative multi-view geometry learning with hybrid-domain enhancement for light field occlusion removal
Jieyu Chen, Ping An 0001, Xinpeng Huang, Chao Yang 0021, Ce Zhu, Sanghoon Lee 0001
Expert Syst. Appl.3
2026 A novel EPI-guided network with progressive fusion for light field reconstruction
Baoshuai Wang, Xinpeng Huang, Ping An 0001
Expert Syst. Appl.3
2026 Enhancing joint human-machine image compression via Chebyshev space modulation
Zhicheng Ma, Ping An 0001, Shipei Wang, Chao Yang 0021, Xinpeng Huang
Multim. Syst.5
2026 Generic feature extraction and compression for human and machine-oriented vision
Kunqiang Huang, Ping An 0001, Chao Yang 0021, Shipei Wang, Xinpeng Huang, Liquan Shen
Signal Process.5
2026 Viewport-Patch Extraction Enhanced 360$^\circ$ Video Quality Assessment
abstract
With the rising adoption of 360$^\circ$video in virtual reality (VR) applications, assessing its perceptual quality remains a challenge due to projection-induced distortions in equirectangular projection (ERP) formats. Traditional sliding-window cropping methods often distort high-latitude content and fail to reflect the actual viewing experience. To address this, we propose a novel viewport patch-based video quality assessment (VQA) method. By sampling view directions on the sphere and applying gnomonic projection, our method extracts undistorted and perceptually consistent viewport patches that preserve both spatial fidelity and full-frame coverage. We further design a two-stream network that jointly models high-frequency distortion and residual information over time, enhanced by squeeze-and-excitation (SE) attention to capture spatial-temporal features. Experiments and analysis show that our method significantly improves the accuracy and reliability of 360$^\circ$VQA, achieving PLCC/SROCC values of 0.9603/0.9628 on the VQA-ODV dataset and 0.9585/0.9400 on the BIT360 dataset, with only 0.22M parameters. Code is available athttps://github.com/yeonhw/VP-VQA.
Chao Yang 0021, Ping An 0001, Xinpeng Huang
IEEE Signal Process. Lett.4
2026 Low-Bitrate Light Field Video Compression Through Key Sequences Encoding and Joint Reconstruction Network
abstract
Light field (LF) videos contain rich spatial, angular, and temporal information, resulting in immense data volumes and posing significant challenges for low-bitrate compression. Existing LF video compression methods focus on modifying the structure of traditional video codecs to encode all LF views, but they are insufficient to achieve low-bitrate compression of LF video. To address these limitations, we propose a low-bitrate LF video compression framework that exploits spatial-angular-temporal correlations through sparse coding and joint reconstruction. On the encoding side, we introduce a content-adaptive prediction structure for sparse key view sequences selection. This structure is adapted to LF video content, leveraging the most similar view as a reference to enhance prediction accuracy and significantly reduce bitrate. On the decoding side, we observe that pixels missing in the current view are often captured in adjacent angular and/or temporal views. As a result, we develop a spatial-angular-temporal based joint reconstruction network that integrates cues across the different domains. This approach supplements missing texture details near occlusion areas and reconstructs high-quality non-key views. Experimental results demonstrate the efficiency of our framework, achieving an average gain of about 60 % in terms of bitrate savings and 2 dB in terms of reconstruction quality compared to the state-of-the-art methods.
Xinpeng Huang, Chao Yang 0021, Mounir Kaaniche, Qiuwen Zhang, Ping An 0001
IEEE Trans. Circuits Syst. Video Technol.2
2026 Capture More, Synthesize Better: Video Frame Interpolation With Larger Receptive Field and Structural Priors
abstract
A large receptive field is crucial for the video frame interpolation (VFI) task. Existing video frame interpolation methods struggle with large motions due to their limited receptive fields. However, simply expanding the receptive field brings two challenges: a substantial computational burden and potential loss of texture details. In this paper, we first propose a novel spatial-temporal global window self-attention mechanism with an enlarged receptive field to enhance motion capture. Furthermore, to reduce the computational complexity introduced by the global window, we design a simple and effective separable fence window decomposition. Meanwhile, to better synthesize high-quality intermediate frames, we propose two complementary frame synthesis strategies. First, from the perspective of receptive field design, we introduce a progressive receptive field focusing module, enabling a smooth transition from global motion modeling to local detail preservation. Second, based on the VFI-specific property and the high structural similarity shared by the adjacent frames, we propose a structure-aware synthesis strategy, which incorporates structural priors to guide the generation of fine details. Subjective and objective experimental results demonstrate that our method effectively captures large motions while synthesizing texture details, outperforming state-of-the-art techniques on various datasets.
Baojun Zhou, Xinpeng Huang, Jieyu Chen, Mounir Kaaniche, Ping An 0001
IEEE Trans. Circuits Syst. Video Technol.2
2026 Volume Feature Aware View-Epipolar Transformers for Generalizable NeRF
abstract
Generalizable NeRF synthesizes novel views of unseen scenes without per-scene training. The view-epipolar transformer has become popular in this field for its ability to produce high-quality views. Existing methods with this architecture rely on the assumption that texture consistency across views can identify object surfaces, with such identification crucial for determining where to reconstruct texture. However, this assumption is not always valid, as different surface positions may share similar texture features, creating ambiguity in surface identification. To handle this ambiguity, this paper introduces 3D volume features into the view-epipolar transformer. These features contain geometric information, which will be a supplement to texture features. By incorporating both texture and geometric cues in consistency measurement, our method mitigates the ambiguity in surface detection. This leads to more accurate surfaces and thus better novel view synthesis. Additionally, we propose a decoupled decoder where volume and texture features are used for density and color prediction respectively. In this way, the two properties can be better predicted without mutual interference. Experiments show improved results over existing transformer-based methods on both real-world and synthetic datasets.
Ping An 0001, Xinpeng Huang, Qiang Wu 0001
IEEE Trans. Vis. Comput. Graph.3
2026 Learning a Domain-Specialized Network for Light Field Spatial-Angular Super-Resolution
abstract
Light field (LF) imaging is inherently constrained by the trade-off between spatial resolution and angular sampling density. To overcome this obstacle, spatial-angular super-resolution (SR) methods have been developed to achieve concurrent enhancement in both dimensions. Traditional spatial-angular SR methods treat spatial and angular SR as separate tasks, resulting in parameter redundancy and error accumulation. While recent end-to-end approaches attempt joint processing, their uniform treatment of these distinct problems overlooks critical domain-specific requirements. To address these challenges, we propose a domain-specialized framework that deploys stage-tailored strategies to satisfy domain-specific demands. Specifically, in the angular SR stage, we introduce a cross-view consistency modulation module that enhances inter-view coherence through long-range dependency modeling of angular features. In the spatial SR stage, we propose a detail-aware state space model to reconstruct fine-grained detail. Finally, we develop a cross-domain integration module that explores spatial-angular correlations by fusing multi-representational features from both domains to foster synergistic optimization. Experimental results on public LF datasets demonstrate substantial improvements over state-of-the-art methods in both qualitative and quantitative comparisons, with approximately 50% fewer model parameters compared to competing methods.
Xinpeng Huang, Deyang Liu, Ping An 0001, Sanghoon Lee 0001
IEEE Trans. Vis. Comput. Graph.2
2025 ContribChain: A Stress-Balanced Blockchain Sharding Protocol with Node Contribution Awareness
Xinpeng Huang, Wanqing Jie, Haofu Yang, Wangjie Qiu, Qinnan Zhang, Huawei Huang, Zehui Xiong, Shaoting Tang, Hongwei Zheng 0003, Zhiming Zheng 0001
INFOCOM1
2025 Agent4Vul: multimodal LLM agents for smart contract vulnerability detection
Wanqing Jie, Wangjie Qiu, Haofu Yang, Muyuan Guo, Xinpeng Huang, Tianyu Lei, Qinnan Zhang, Hongwei Zheng 0003, Zhiming Zheng 0001
Sci. China Inf. Sci.5
2025 Layered and scalable image coding with semantic features for human and machine
Jiao Wei, Ping An 0001, Shipei Wang, Kunqiang Huang, Chao Yang 0021, Xinpeng Huang
Eng. Appl. Artif. Intell.6
2025 SADiff: Structure-aware diffusion model for enhancing prostate multiphoton microscopy imaging
Maoye Huang, Xinpeng Huang, Haoyi Fan, Xiaoqin Zhu
Expert Syst. Appl.2
2025 Disparity Enhancement-Based Light Field Angular Super-Resolution
abstract
The depth-dependent light field (LF) reconstruction is a prevalent solution for large disparity LFs, which first estimates disparity maps and subsequently interpolates the content of target views by warping input views. However, replication errors often occur in edge regions of objects owing to inappropriate sampling positions caused by occlusion during disparity-based warping. Thus, we propose a disparity enhancement network that utilizes morphological filtering to address this distortion, which can adaptively modify disparity values in edge regions to obtain proper sampling positions. In addition, we develop an effective detail recovery network to mitigate interpolation errors introduced by inaccurate disparity estimation or warping operations. Experiments demonstrate that our approach significantly surpasses current state-of-the-art methods in large disparity LFs.
Dongjun Cai, Xinpeng Huang, Ping An 0001
IEEE Signal Process. Lett.3
2025 Mask-Aware Light Field De-Occlusion With Gated Feature Aggregation and Texture-Semantic Attention
abstract
A light field image records rich information of a scene from multiple views, thereby providing complementary information for occlusion removal. However, current occlusion removal methods have several issues: 1) inefficient exploitation of spatial and angular complementary information among views; 2) indistinguishable treatment of pixels from foreground occlusion and background; and 3) insufficient exploration of spatial detail supplementation. Therefore, in this article, we propose a mask-aware de-occlusion network (MANet). Specifically, MANet is a joint training network that integrates the occlusion mask predictor (OMP) and the occlusion remover (OR). First, OMP is proposed to provide the location of occluded regions for OR, as the occlusion removal task is ill-posed without occluded region localization. In OR, we introduce gated spatial-angular feature aggregation, which uses a soft gating mechanism to focus on spatial-angular interaction features in non-occluded regions, extracting effective aggregated features specific to the de-occlusion. Then, we design a complementary strategy to fully utilize spatial-angular information among views. Finally, we propose texture-semantic attention to improve the performance of detail generation. Experimental results demonstrate the superiority of MANet, with substantial improvements in both PSNR and SSIM metrics. Moreover, MANet stands out with an efficient parameter count of 2.4 M, making it a promising solution for real-world applications in public safety and security surveillance.
Jieyu Chen, Ping An 0001, Xinpeng Huang, Chao Yang 0021, Liquan Shen
IEEE Trans. Multim.3
2025 Feature Quality Assessment: A Database and A Lightweight Objective Method
abstract
In the era of Artificial Intelligence, visual data gathered by edge devices could be primarily utilized for machine vision tasks. The prominent coding frameworks accomplish this by extracting and compressing features extracted from input data. As such, the quality of these features is vital, as they reflect the performance of the coding framework. However, much less work has been dedicated to quality assessment on features, impeding the optimization of the coding system. In this work, we pioneer to explore the feature quality assessment by creating a novel database tailored for features, with the quality ground-truth for each feature. Then, we propose a lightweight feature quality assessment method, called Lightweight Feature Quality Assessment (LFQA). We analyze the feature characteristics from the perspective of spatial and channel thoroughly, and the framework of LFQA is designed based on the analysis results. Experimental results demonstrate that LFQA accurately evaluates the quality of features, reaching a notable Spearman Rank-Order Correlation Coefficient of 85.38%, and exhibits competitive performance in improving the performance of video coding for machine system. Furthermore, LFQA has fewer model parameters and faster inference speed, ensuring a wide range of promising applications.
Shipei Wang, Ping An 0001, Chao Yang 0021, Gongyang Li, Xinpeng Huang, Shiqi Wang 0001
IEEE Trans. Multim.5
2025 Dual-Guided Video Frame Interpolation With Spatial-Temporal Global Attention
abstract
Video frame interpolation technology improves visual experience with the development of deep learning. However, capturing large motions while synthesizing fine texture details remains a challenging task. Regarding large motion scenarios, some pioneering Transformer-based methods primarily rely on local attention, which does not fully leverage the global receptive field advantage. To address this issue, this paper proposes to further broaden the receptive field of the Transformer to capture more correlations in the video frame interpolation task. Specifically, we propose a global self-attention mechanism in the form of spatial-temporal separation. Regarding texture details, since roughly enlarging the receptive field results in the loss of details, we propose to use large motion information in both feature and pixel spaces as a dual-guided prior to enhance detail synthesis. The separable attention mechanism and the straightforward frame synthesis design significantly enhance the resource efficiency of our model. Extensive experiments show that our method achieves state-of-the-art performance, effectively capturing large motions and preserving texture details.
Baojun Zhou, Xinpeng Huang, Gongyang Li, Chao Yang 0021, Liquan Shen, Ping An 0001
IEEE Trans. Multim.2
2024 Low Bitrate Light Field Video Compression with Two-step Refinement Reconstruction
abstract
Light Fields (LFs) are characterized by high dimension, complex structure and large amount of data. Therefore, the efficient compression of LF videos faces challenges. Existing methods use the traditional multi-view video encoder to compress LF videos, but the bitrates are still high. To this end, we propose to only encode sparse key view sequences in LF video for low bitrates. The proposed similarity-based prediction structure fully exploits the spatial-angular-temporal correlations in LF videos. At the decoder side, we propose to reconstruct the uncoded non-key view sequences by a two-step refinement reconstruction network. To avoid error propagation, the proposed reconstruction network firstly refines the coded key view sequences by reducing the compression artifacts. With regard to the distortion caused by warping, the network further refines the texture details of the synthesized non-key view sequences. The experiment results prove that the proposed algorithm performs superior at low bitrates.
Xinpeng Huang
ICME2
2024 Light Field Stitching Via Mesh Deformation Alignment and Low-Rank-Based Fusion
Xinpeng Huang
PRCV (9)4
2024 Adaptive Threshold Mask Prediction and Occlusion-aware Convolution for Foreground Occlusions in Light Fields
abstract
The performance of existing de-occlusion methods is limited mainly due to inaccurate foreground mask prediction, interference from occlusion information in feature extractors, and insufficient utilization of sub-pixel information between views. Therefore, in this paper, we propose an efficient light field image de-occlusion method to improve the performance of occlusion localization and occlusion removal. First, we design an adaptive threshold mask prediction branch, which fully utilizes the spatial and angular information of light field images and incorporates mask binarization into the network for joint optimization. Then, we propose Occlusion-aware Convolution, which can more efficiently extract the joint spatial-angular features of the occluded light field image. Additionally, we design a sub-pixel complement strategy. This strategy fully utilizes the sub-pixel information between the views and supplements it into the view to be deoccluded. Experimental results demonstrate that our method achieves superior performance on both real-world and synthetic light field datasets.
Jieyu Chen, Ping An 0001, Xinpeng Huang, Chao Yang 0021
VCIP3
2024 TSARN: A Joint Temporal-Spatial-Angular Reconstruction Network for Light Field Lenslet Video Compression
abstract
Light Field (LF) lenslet videos are capable of capturing the aggregation of light rays in dynamic scenes, providing a more immersive experience. However, the immense data volume of LF lenslet videos poses significant challenges for storage and transmission. Existing methods that utilize traditional video codecs to encode all views fail to meet the demands for efficient compression. This paper proposes a compression framework based on down-sampling, which focuses on designing an efficient reconstruction method at the decoder side. In LF lenslet videos, pixels that are invisible in the current view may be visible in adjacent views angularly and/or in adjacent frames temporally. Therefore, we propose a joint temporal-spatial-angular reconstruction network. In this network, Spatial-Angular Convolutional Module utilizes different forms of convolutions to fully extract spatial-angular features for synthesizing LF views initially; Deformable Convolutional Temporal Fusion Module employs deformable convolutions to align information from different frames and aggregate temporal features; View Refinement Module is finally adopted to further refine the features and reconstruct high-quality LF lenslet videos. Experimental results demonstrate that our proposed LF lenslet video compression framework achieves superior performance, significantly reducing the bitstream.
Xinpeng Huang, Yongjie Lu
VCIP2
2024 Low-Rate Feature Compression for Humans and Machines with Dual Aggregation Attention
abstract
The Collaborative Intelligence (CI) framework offers innovative approaches for deploying Deep Neural Networks (DNNs). However, the limitations of communication resources require minimizing the transmission of bits between the edge and cloud devices to meet the requirements of both machine recognition and human perception. Previous research has demonstrated that the transmission of intermediate layer features of the vision backbone can achieve superior performance in machine vision tasks at very low bit rates without consuming additional bits. Nonetheless, the reduced bit rate poses challenges for image reconstruction. We propose a CI framework that compresses the intermediate features of the Swin Transformer and utilizes a Feature Recovery Module (FRM) to restore crucial information for image reconstruction, thereby satisfying both machine and human visual tasks at low bit rates. Additionally, we introduce a Residual Dual-attention Aggregation Block (RDAB) that exploits both local and global information for effective compression and reconstruction. We conducted experiments on the CUB_200_2011 dataset. The results demonstrate that the proposed method delivers superior performance at low-rate scenarios.
Ruixi Ma, Ping An 0001, Shipei Wang, Xinpeng Huang, Chao Yang 0021
VCIP4
2024 Content adaptive spatial-temporal rescaling for video coding optimization
Chao Yang 0021, Siqian Qin, Ping An 0001, Xinpeng Huang, Liquan Shen
Expert Syst. Appl.4
2024 Learning-based CU partition prediction for fast panoramic video intra coding
Chao Yang 0021, Ping An 0001, Xinpeng Huang, Liquan Shen
Expert Syst. Appl.4
2024 STSIC: Swin-transformer-based scalable image coding for human and machine
Shipei Wang, Ping An 0001, Chao Yang 0021, Kunqiang Huang, Xinpeng Huang
J. Vis. Commun. Image Represent.5
2024 Scalable image coding with enhancement features for human and machine
Ping An 0001, Chao Yang 0021, Xinpeng Huang
Multim. Syst.4
2024 360° video quality assessment based on saliency-guided viewport extraction
Fanxi Yang, Chao Yang 0021, Ping An 0001, Xinpeng Huang
Multim. Syst.4
2024 Light Field Salient Object Detection With Sparse Views via Complementary and Discriminative Interaction Network
abstract
4D light field data record the scene from multiple views, thus implicitly providing beneficial depth cue for salient object detection in challenging scenes. Existing light field salient object detection (LF SOD) methods usually use a large number of views to improve the detection accuracy. However, using so many views for LF SOD brings difficulties to its practical applications. Considering that adjacent views in a light field are actually with very similar contents, in this work, we propose defining a more efficient pattern of input views, i. e., key sparse views, and design a network to effectively explore the depth cue from sparse views for LF SOD. Specifically, we firstly introduce a low rank-based statistical analysis to the existing LF SOD datasets, which allows us to conclude a fixed yet universal pattern for our key sparse views, including the number and positions of views. These views maintain the sufficient depth cue, but greatly lower the number of views to be captured and processed, facilitating practical applications. Then, we propose an effective solution with a key Complementary and Discriminative Interaction Module (CDIM) for LF SOD from key sparse views, named CDINet. The CDINet follows a two-stream structure to extract the depth cue from the light field stream (i. e., sparse views) and the appearance cue from the RGB stream (i. e., center view), generating features and initial saliency maps for each stream. The CDIM is tailored for inter-stream interaction of both these features and saliency maps, using the depth cue to complement the missing salient regions in RGB stream and discriminate the background distraction, to enhance the final saliency map further. Extensive experiments on three LF multi-view datasets demonstrate that our CDINet not only outperforms the state-of-the-art 2D methods, but also achieves competitive performance as compared with the state-of-the-art 3D and 4D methods. The code and results of our method are available athttps://github.com/GilbertRC/LFSOD-CDINet.
Gongyang Li, Ping An 0001, Zhi Liu 0003, Xinpeng Huang, Qiang Wu 0001
IEEE Trans. Circuits Syst. Video Technol.5
2023 Intermediate deep feature coding for human-machine vision collaboration
Weiqian Wang, Ping An 0001, Xinpeng Huang, Kunqiang Huang, Chao Yang 0021
J. Vis. Commun. Image Represent.3
2022 An improved lossless image compression algorithm based on Huffman coding
Ping An 0001, Xinpeng Huang
Multim. Tools Appl.4
2022 Energy-driven reference selection for hierarchical light field compression
Xinpeng Huang, Ping An 0001, Deyang Liu
Signal Process. Image Commun.1
2022 Light field occlusion removal network via foreground location and background recovery
Shiao Zhang, Ping An 0001, Xinpeng Huang, Chao Yang 0021
Signal Process. Image Commun.4
2022 Low Bitrate Light Field Compression With Geometry and Content Consistency
abstract
Light field imaging can simultaneously record the position and direction information of light rays; thus, digital refocusing and full depth-of-field extension — functions that are inaccessible for conventional images — can be achieved using the structural consistency of light field data. To meet the challenges of limited bandwidth and storage, such vast numbers of light field data must be compressed to a low bitrate. However, current compression solutions ignore the intrinsic consistency of light fields in pursuit of a low bitrate, thereby leading to the loss of light field capabilities. To solve this issue, this work focuses on structural consistency to achieve efficient light field compression with a low bitrate. The proposed light field compression method encodes the sparsely selected sub-aperture images (SAIs) and the disparity maps corresponding to the unselected SAIs. From the perspective of geometry consistency, the consistency of the initially estimated disparity maps is improved by using a color-guided refinement algorithm, thereby reducing the bitrate of the disparity maps. From the perspective of content consistency, the consistency of the SAI-transformed pseudo sequence is improved by the proposed content-similarity-based arrangement algorithm along with a specific prediction structure; thereby, the bitrate of the sparsely selected SAIs is reduced. The experimental results show that the proposed compression method can reduce the total bitrate while preserving good structural consistency.
Xinpeng Huang, Ping An 0001, Deyang Liu, Liquan Shen
IEEE Trans. Multim.1
2022 Objective Quality Assessment of Lenslet Light Field Image Based on Focus Stack
abstract
The large amount of complex scene information recorded by light field imaging has the potential for immersive media applications. Compression and reconstruction algorithms are crucial for the transmission, storage, and display of such massive data. Most of the existing quality evaluation indexes do not effectively account for light field characteristics. To accurately evaluate the distortions caused by compression and reconstruction algorithms, it is necessary to construct an image evaluation index that reflects the angular-spatial characteristics of the light field. This work proposes a full-reference light field image quality evaluation index that attempts to extract less information from the focus stack to accurately evaluate the entire light field quality. The proposed framework includes three specific steps. First, we construct a key refocused image extraction framework by the maximal spatial information contrast and the minimal angular information variation. Specifically, the gradient and phase congruency operators are used in the extraction framework. Second, a novel light field quality evaluation index is built based on the angular-spatial characteristics of the key refocused images. In detail, the features used in the key refocused image extraction framework and the chrominance feature are combined to construct the union feature. Third, the similarity of the union feature is pooled by the relevant visual saliency map to obtain the predicted score. Finally, the overall quality of the light field is measured by applying the proposed index to the key refocused images. The high efficiency and precision of the proposed method are shown by extensive comparison experiments.
Chunli Meng, Ping An 0001, Xinpeng Huang, Chao Yang 0021, Liquan Shen
IEEE Trans. Multim.3
2021 Multi-Models Fusion for Light Field Angular Super-Resolution
abstract
Light field (LF) imaging has received increasing attention due to its richer interpretation of the scene. However, an inherent spatial-angular trade-off exists in LF that prevents LF from practical applications. Consequently, how to break such a trade-off has become one of the main challenges in sparsely sampled LF reconstruction. LF super-resolution (SR) can provide an opportunity to solve this issue, but most methods exploit only one form of LF, thereby leading to much loss of information. We believe that different LF forms can compensate each other to obtain higher gains via fusion strategy. In this paper, therefore, we propose a multi-models fusion for LF SR in angular domain. Cascading models which are trained by different LF forms can fully exploit rich LF information. Experimental results demonstrate that our method is effective and achieves a comparable result against state-of-the-art techniques.
Fengyin Cao, Ping An 0001, Xinpeng Huang, Chao Yang 0021, Qiang Wu 0001
ICASSP3
2021 Fast And Accurate Homography Estimation Using Extendable Compression Network
abstract
Fast and accurate homography estimation between images is crucial for relative pose estimation in autonomous exploration. Recently, learning-based methods have been proposed to use semantic information to solve challenging cases like large displacements, dynamic scenes, and illumination changes, where traditional methods may degrade. However, most existing methods have large model sizes and low inference speed, which make them infeasible in terminal devices and real-time scenarios. In this paper, we build a basic network based on the ShuffleNetV2 compressed units, which can extremely accelerate the homography estimation process. To further deal with the large displacements, we extend the basic network to a multiscale weight-shared form to additionally process the half-scale input. In the case of sufficient computational resources, this basic network can also be extended to a recurrent coarse-to-fine form to achieve the most accurate results. Experimental results show that our extendable networks can well balance the accuracy and inference speed, and the sizes of all models are less than 9MB.
Zhixiang You, Xinpeng Huang
ICIP5
2021 View synthesis-based light field image compression using a generative adversarial network
Deyang Liu, Xinpeng Huang, Wenfa Zhan, Liefu Ai, Shulin Cheng
Inf. Sci.2
2021 Learning from EPI-Volume-Stack for Light Field image angular super-resolution
Deyang Liu, Qiang Wu 0001, Yan Huang 0023, Xinpeng Huang, Ping An 0001
Signal Process. Image Commun.4
2020 Light Field Compression Using Global Multiplane Representation and Two-Step Prediction
abstract
Due to its spatio-angular structure, light field image allows for a wealth of post-processing techniques like digital refocusing and depth estimation. In order to compress the data of the two domains, the current proposal intends to embed the disparity-based view synthesis method into the decoder. However, predicting each view separately or in local groups means bringing more computational burden to the decoder and destroying the light field structure. Since disparity contains the relationship between all light rays in the light field, the proposed solution is to predict a disparity-based global representation as the first step. In the second step, all the views can be predicted easily based on this representation. In this letter, we use the recently proposed multiplane as the form of this global representation. The experimental results show the effectiveness of the proposed solution, and the better RD performance compared to other schemes especially under low bitrates.
Ping An 0001, Xinpeng Huang, Chao Yang 0021, Deyang Liu, Qiang Wu 0001
IEEE Signal Process. Lett.3
2020 Full Reference Light Field Image Quality Evaluation Based on Angular-Spatial Characteristic
abstract
The quality evaluation is an indispensable link in light field (LF) image processing. Most of existing LF objective evaluation indexes do not make effective use of the angular characteristic of LF, so the evaluation results are unsatisfactory. In this letter, the quality evaluation of LF image is constructed based on human visual system (HVS) and LF angular-spatial characteristics. Based on the fact that HVS has different sensitivity to different parallaxes, we assume that LF image quality perceived by human eyes has the optimal parallax range. A dual-fan filter is used to constrain the parallax range. Then, the overall quality of the LF is represented by combining the spatial and angular quality, which performed from the central sub-aperture image and the focus stack, respectively. In addition, because the difference of Gaussian (DoG) operator can simulate the process of extracting texture structure by human eyes. The structural similarity of DoG texture feature is utilized in the spatial quality evaluation. Extensive comparison experiments show that the proposed method is more consistent with the characteristics of LF.
Chunli Meng, Ping An 0001, Xinpeng Huang, Chao Yang 0021, Deyang Liu
IEEE Signal Process. Lett.3
2020 Content-Based Light Field Image Compression Method With Gaussian Process Regression
abstract
Light field (LF) imaging enables new possibilities for digital imaging, such as digital refocusing, changing of focus plane, changing of viewpoint, scene-depth estimation, and 3D scene reconstruction, by capturing both spatial and angular information of light rays. However, one main problem in dealing with LF data is its sheer volume. In this context, efficient compression methods are needed for such a particular type of content. In this paper, we propose a content-based LF image-compression method with Gaussian process regression to improve the compression efficiency and accelerate the prediction procedure. First, the LF image is fed to the intra-frame codec of HEVC. In the prediction procedure, the prediction units (PUs) are classified as non-homogenous texture units, homogenous texture units, and visually flat units, based on the content property of the LF image. For each category, we design a corresponding Gaussian process regression (GPR)-based prediction method. Moreover, we propose a classification mechanism to exactly decide to which category the current PU belongs, so as to adjust the trade-off between the computational burden and the LF image coding efficiency. Experimental results demonstrate that the proposed LF image compression method is superior to several other state-of-the-art compression methods in terms of different quality metrics. Furthermore, the proposed method can also achieve a good visual quality of views rendered from decoded LF contents.
Deyang Liu, Ping An 0001, Wenfa Zhan, Xinpeng Huang, Ali Abdullah Yahya
IEEE Trans. Multim.5
2019 Objective Quality Assessment for Light Field Based on Refocus Characteristic
Chunli Meng, Ping An 0001, Xinpeng Huang, Chao Yang 0021
ICIG (3)3
2019 Modified Baseline for Light Field Stitching
abstract
In traditional 2D image stitching, the baseline method usually means global homography via Direct Linear Transformation (DLT) on inliers. In this paper, a modified baseline method for light field (LF) stitching is proposed to stitch two LFs. The depth map and the center sub-aperture image (SAI) are used to filter the feature points of the entire LF. The global 4D homography is then calculated by DLT to align all SAIs corresponding to the same angular domain coordinates of two LFs. Finally, the improved Markov Random Field (MRF) energy considering the global LF is used to find the seam of 2D SAIs instead of computational 4D graph cut. Experimental results show that the proposed method can effectively stitch the 4D LFs, and preserve the consistency of the angular and spatial domains of the stitched LF compared with implementing 2D image stitching to the corresponding SAIs. Moreover, the method proposed in this paper can easily extend all advanced 2D image stitching methods to 4D LF, so that the acquired LF can have larger field of view and wider applications.
Ping An 0001, Xinpeng Huang, Chunli Meng, Qiang Wu 0001
VCIP3
2018 View Synthesis for Light Field Coding Using Depth Estimation
abstract
Light Field (LF) image captured by plenoptic camera can record richer scenario information from our world. But a huge number of Sub-Aperture Images (SAIs) from LF image results in great challenges for coding LFI. Therefore, we choose a subset of SAIs as multi-view video, and encode them with their corresponding depth maps using Multi-view Video plus Depth (MVD) structure. Due to lack of depth maps for SAIs, we propose a cost function to determine the depth map for each SAI preliminarily based on horizontal and vertical Epipolar Plane Images (EPI), respectively. Then an SAI-guided depth enhancement algorithm is designed to optimize the estimated depth maps. Since those unselected SAIs have not been encoded but have been synthesized using the specific texture image and depth map, our LF image coding method can naturally achieve bitrates reduction dramatically with a good performance and outperform other algorithms significantly.
Xinpeng Huang, Ping An 0001, Liang Shan 0004, Liquan Shen
ICME1
2018 Light Field Image Sparse Coding via CNN-Based EPI Super-Resolution
abstract
This paper proposes a novel light field (LF) image compression scheme by super resolving the epipolar plane image (EPI) via convolutional neural network (CNN). In the scheme, we first decompose the LF image into sub-aperture images (SAIs), and only one quarter of them are compressed on the encoding side to reduce the bitrate. On the decoding side, we use these selected SAIs to reconstruct the entire LF by taking advantage of the special structure of EPI. The low-resolution EPIs generated from the sparse SAIs are super resolved by using deep residual network and the output high-resolution EPIs are used to rebuild the dense SAIs. Experimental results show the superior performance of our scheme, which achieve 1.46 dB quality improvement and 35.85 percent bit rate reduction on average compared with the typical pseudo-sequence-based coding method.
Jinbo Zhao, Ping An 0001, Xinpeng Huang, Liang Shan 0004
VCIP3