VLDB 2026 Research / reviewers in the wild / expert
Deyang Liu
dblp:63/8755
· DBLP profile ↗
44ranked-venue papers
18as first author
33since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 36 · 15 first-author · 26 since 2021Artificial intelligence and machine learning · 6 · 2 first-author · 5 since 2021Computer networks · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A blockchain-enhanced cross-domain handover authentication scheme for Vehicular Ad-hoc Networks
Wenming Wang 0001, Deyang Liu, Zhuofei Wu, Haiping Huang |
Ad Hoc Networks | 3 |
| 2026 | Low-light light field image enhancement based on illumination-guided implicit gradient representation
Deyang Liu, Xiaofei Zhou 0003, Ping An 0001, Caifeng Shan, Hongbin Zha |
Neurocomputing | 1 |
| 2026 | SEBA: A Secure and Efficient Blockchain-Based Cross-Domain Authentication Scheme for Vehicular NetworksabstractWith the rapid growth of connected vehicles and the increasing demand for low-latency and reliable communication, traditional vehicular networks face severe security challenges. Existing cross-domain authentication schemes often suffer from complex certificate management, key leakage, and inefficient trust coordination. In this paper, we propose a Secure and Efficient Blockchain-based Cross-Domain Authentication (SEBA) scheme for vehicular networks that integrates blockchain with physical unclonable functions (PUFs) to establish a robust trust management framework. The proposed scheme adopts a hybrid on-chain/off-chain architecture, where blockchain-based smart contracts provide tamper-resistant credential storage, while time-critical authentication operations are executed off-chain to improve system throughput and response latency. By incorporating PUF technology, SEBA enables reliable binding between physical devices and cryptographic identities, effectively resisting impersonation and cloning attacks. In addition, lightweight cryptographic primitives are employed to further reduce computation and communication overhead. The security of the proposed scheme is formally verified using ProVerif and further analyzed through theoretical security analysis. Performance evaluation is conducted using a simulation framework that models vehicular communication latency, cryptographic operation costs, and blockchain transaction confirmation delays. Experimental results under varying vehicle loads demonstrate that SEBA significantly outperforms existing schemes, achieving up to 83% higher throughput and 38% lower authentication latency. Overall, SEBA provides a privacy-preserving and low-latency authentication framework suitable for large-scale and highly dynamic vehicular networks. Wenming Wang 0001, Deyang Liu, Zhiquan Liu 0001, Haiping Huang |
IEEE Internet Things J. | 3 |
| 2026 | Dense multiscale inference network for lightweight salient object detection of strip steel surface defects
Yihan Qiu, Xiaofei Zhou 0003, Yong Wu 0007, Bin Wan, Juting Miu, Zhangping Chen, Deyang Liu |
Pattern Recognit. Lett. | 7 |
| 2026 | Global and local collaborative learning for no-reference omnidirectional image quality assessment
Deyang Liu, Lifei Wan, Xiaofei Zhou 0003, Caifeng Shan |
Signal Process. Image Commun. | 1 |
| 2026 | Degradation-Aware Blind Light-Field Image Quality Assessment With Linear AttentionabstractBlind Light-Field Image Quality Assessment (LFIQA) is challenging, as degradations are microlens-dependent and spatially non-uniform, while perceptual quality relies on both spatial fidelity and angular consistency. However, many existing methods either assume globally stationary distortions or adopt global pooling or self-attention, which can be biased by locally corrupted lenslets and become computationally prohibitive when modeling long-range spatial-angular dependencies. Therefore, in this letter, we present a degradation-aware framework that first predicts a microlens reliability map to make quality inference robust to spatially non-uniform, lenslet-varying corruption. It then extracts spatial and angular features, applies a shared Receptance Weighted Key Value (RWKV) module for linear-time long-range context, and fuses them to predict perceptual quality. Experiments show a higher correlation with subjective ratings with competitive efficiency. Youzhi Zhang 0004, Jianyu Qian, Deyang Liu, Xiaofei Zhou 0003, Hongbin Zha, Caifeng Shan |
IEEE Signal Process. Lett. | 3 |
| 2026 | Learning Implicit and Detail-Enhanced Network for Light Field Image Spatial-Angular Super-ResolutionabstractLight field (LF) imaging holds immense promise for applications such as post-capture refocusing and virtual reality. However, its inherent spatial-angular trade-off significantly limits both spatial and angular resolution, restricting its practicality in real-world scenarios. To address these limitations, spatial-angular super-resolution methods have been proposed to simultaneously enhance both dimensions. Yet, existing methods struggle to fully exploit the intertwined spatial-angular correlations and fail to effectively handle sparsely sampled LFs with low spatial resolution, often leading to cumulative errors during reconstruction. In this paper, we propose an Implicit and Detail-Enhanced Network (IDNet) to overcome these challenges. Our IDNet employs 3D convolution for the joint extraction of spatial and angular information, leveraging their interdependencies for more effective LF reconstruction. Additionally, we introduce an implicit detail restoration module that enhances features while encoding positional information to refine fine details. To overcome the limitations of sparse spatial and angular information on high-detail reconstruction and angular consistency in low-resolution LFs, we design a multi-representation enhancement block. This block enhances features by learning pixel differences across multiple directions in diverse representations, effectively capturing intricate details and complex correlations. Thanks to these designs, our IDNet reconstructs novel views with finer details, effectively learns occlusion relationships, and ensures geometric consistency. Experimental results on benchmark datasets demonstrate its superior quantitative and qualitative performance. The code is publicly available at https://github.com/ldyorchid/IDNet. Deyang Liu, Shizheng Li, Xiaofei Zhou 0003, Zeyu Xiao 0002, Caifeng Shan |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2026 | Probabilistic-Based Learning for Joint Light Field Image Compression and Enhancement Under Low-Light ConditionsabstractLight field (LF) imaging has attracted increasing research interest in challenging illumination conditions due to its ability to provide rich spatial and angular cues. However, such data present dual challenges: 1) the inherent multi-view structure introduces substantial data redundancy, creating high demands for efficient compression; 2) the insufficient illumination leads to severe quality degradation, which weakens inter-view consistency and visual perception. To address these coupled factors, we propose a Probabilistic-based learning for joint LF image compression and enhancement under low-light conditions (PrL-LFCE). The framework unifies structure-aware compression and feature enhancement mechanisms by introducing learnable probabilistic modeling into both feature coupling and latent distribution estimation to adaptively handle the uncertainty induced by illumination degradation and compression-related information loss. Specifically, we design a probability-based multi-directional feature coupling module that dynamically balances structural preservation and redundancy reduction across multiple directionally arranged sub-aperture images. Moreover, we introduce a swin-gated enhancement module that suppresses noise and highlights structurally salient regions in compression-aware feature representations through attention-guided gating. Extensive experiments show that PrL-LFCE consistently outperforms state-of-the-art methods, achieving at least 34.86% bitrate savings while maintaining excellent visual quality, demonstrating a strong joint compression and enhancement capability. Deyang Liu, Jimin Wang, Mounir Kaaniche, Xiaofei Zhou 0003, Gangyi Jiang, Caifeng Shan |
IEEE Trans. Image Process. | 1 |
| 2026 | Scale-Invariant Feature Matching Network for V-D-T Few-Shot Semantic SegmentationabstractMulti-modal few-shot semantic segmentation (FSS) aims to perform dense prediction from multiple modality images including visible image, depth image, and thermal image with a few annotated samples. However, some efforts treat the three modality information equally, where they don't incorporate the inherent differences among multiple modalities. Besides, the objects vary in size greatly, and the cutting-edge matching paradigms fail to establish an effective support-query connection. Therefore, we propose a novel scale-invariant feature matching network (i.e., SFM-Net), which consists of an encoder, a feature matching block, a feature elevation block, and a decoder, to conduct visible-depth-thermal (V-D-T) few-shot semantic segmentation. Firstly, in the encoder part, after the extraction of multi-level initial features, we fuse each level's RGB feature and thermal feature, yielding the support features and the query features. Secondly, in the feature matching block, a pixel-to-patch cross-attention (PTPCA) module is deployed to explore the correlation between each level's support feature and the query feature, where the pixel-to-patch pooling (PTP-pool) units are designed to build scale-invariant relationships, generating the coarse mask for the query image. Thirdly, in the feature elevation block, we employ the prior-related fusion (PF) module to integrate the depth image with a coarse mask via the cross-attention mechanism, yielding the enhanced coarse prediction result, which is further aggregated in a bottom-up way. Finally, in the decoder, we deploy a reverse attention (RA) unit to gradually explore the complementarity between object internal regions and spatial details, and further generate the final segmentation results via conventional convolution layers. Extensive experiments are conducted on the VDT-2048- $5^{i}$ dataset, and the results show that our model outperforms the state-of-the-art methods with a large margin. Xiaofei Zhou 0003, Deyang Liu, Jiyong Zhang 0001, Runmin Cong |
IEEE Trans. Image Process. | 4 |
| 2026 | Few-Shot Strip Steel Surface Defect Segmentation via Pre-Trained Variational Auto-Encoder-Based Latent Gaussian Process RegressionabstractRecently, few-shot strip steel surface defect segmentation has received more and more concerns. However, the existing few-shot segmentation methods usually adopt the frozen encoder, which is pre-trained on the classification task and can only provide class-related knowledge. Therefore, we propose a novel method, namely pre-trained variational auto-encoder based latent gaussian process regression (LGPR), to conduct few-shot strip steel surface defect segmentation. Firstly, different from previous methods, the frozen Variational Auto-Encoder (VAE) based encoder and decoder, which are pre-trained by using the pixel-level self-supervised task (i.e., image reconstruction), can provide rich image-related knowledge. This ensures the effective characterization of defect regions. Secondly, by deploying a gaussian process regression in the latent feature space generated by the VAE-based encoder, pixel-level correlation between support features and query features can be efficiently built. This operation is non-parametric and doesn't bring any training overhead. Besides, we deploy transformer-based projectors to dig long-range contextual cues of support and query features. Extensive experiments are performed on two public datasets, and the experimental results clearly show that our model consistently outperforms the state-of-the-art models with a large margin. Both the codes and results are publicly available at https://github.com/Hlao-hub/LGPR. Xiaofei Zhou 0003, Gongyang Li, Deyang Liu, Qingshan She, Xiaobin Xu 0002, Runmin Cong |
IEEE Trans. Image Process. | 4 |
| 2026 | Learning a Domain-Specialized Network for Light Field Spatial-Angular Super-ResolutionabstractLight field (LF) imaging is inherently constrained by the trade-off between spatial resolution and angular sampling density. To overcome this obstacle, spatial-angular super-resolution (SR) methods have been developed to achieve concurrent enhancement in both dimensions. Traditional spatial-angular SR methods treat spatial and angular SR as separate tasks, resulting in parameter redundancy and error accumulation. While recent end-to-end approaches attempt joint processing, their uniform treatment of these distinct problems overlooks critical domain-specific requirements. To address these challenges, we propose a domain-specialized framework that deploys stage-tailored strategies to satisfy domain-specific demands. Specifically, in the angular SR stage, we introduce a cross-view consistency modulation module that enhances inter-view coherence through long-range dependency modeling of angular features. In the spatial SR stage, we propose a detail-aware state space model to reconstruct fine-grained detail. Finally, we develop a cross-domain integration module that explores spatial-angular correlations by fusing multi-representational features from both domains to foster synergistic optimization. Experimental results on public LF datasets demonstrate substantial improvements over state-of-the-art methods in both qualitative and quantitative comparisons, with approximately 50% fewer model parameters compared to competing methods. Xinpeng Huang, Deyang Liu, Ping An 0001, Sanghoon Lee 0001 |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2025 | PSReg: Prior-guided Sparse Mixture of Experts for Point Cloud RegistrationabstractThe discriminative feature is crucial for point cloud registration. Recent methods improve the feature discriminative by distinguishing between non-overlapping and overlapping region points. However, they still face challenges in distinguishing the ambiguous structures in the overlapping regions. Therefore, the ambiguous features they extracted resulted in a significant number of outlier matches from overlapping regions. To solve this problem, we propose a prior-guided SMoE-based registration method to improve the feature distinctiveness by dispatching the potential correspondences to the same experts. Specifically, we propose a prior-guided SMoE module by fusing prior overlap and potential correspondence embeddings for routing, assigning tokens to the most suitable experts for processing. In addition, we propose a registration framework by a specific combination of Transformer layer and prior-guided SMoE module. The proposed method not only pays attention to the importance of locating the overlapping areas of point clouds, but also commits to finding more accurate correspondences in overlapping areas. Our extensive experiments demonstrate the effectiveness of our method, achieving state-of-the-art registration recall (95.7%/79.3%) on the 3DMatch/3DLoMatch benchmark. Moreover, we also test the performance on ModelNet40 and demonstrate excellent performance. Xiaoshui Huang, Zhou Huang 0006, Yifan Zuo 0001, Yongshun Gong, Chengdong Zhang, Deyang Liu, Yuming Fang 0001 |
AAAI | 6 |
| 2025 | Structure-Aware and Implicit Restoration for Low-Light Light Field Image CompressionabstractWith the wide application of imaging technologies in low-light scenarios such as night surveillance and astronomical observation, low-light light field (LF) image has become increasingly critical. However, such imaging data face dual challenges. First, the multi-view structure inherently requires high storage and transmission bandwidth, as well as efficient compression. Second, under limited illumination conditions, the reconstructed images exhibit significant noise, low contrast, and detail loss, which severely compromise enhancement performance. Current approaches typically process compression and enhancement sequentially, which suffers from two critical drawbacks: (a) compression artifacts tend to eliminate high-frequency components essential for subsequent enhancement, and (b) existing enhancement algorithms demonstrate limited robustness when handling heavily compressed images. To bridge this gap, we propose a joint compression-enhancement framework that integrates epipolar plane image (EPI) structure-aware modeling into the compression process and incorporates feature-guided restoration during the decoding stage, thereby optimizing the bitrate-distortion-perception trade-off. Specifically, the EPI dual-axis attention (EDA) module is proposed to exploit the horizontal and vertical EPI consistency structures extracted along sub-aperture image (SAI) sequences. Furthermore, the Decoder-side Implicit Detail Restoration (D-IDR) module is designed to restore detailed features. This restoration is enabled by continuous implicit mapping within the compressed latent space, which integrates coordinate-based positional encoding and frequency-domain feature modeling. Experiments on public low-light LF datasets demonstrate that our method outperforms state-of-the-art (SOTA) methods, achieving up to 37.40 % bit-rate savings while maintaining superior structural preservation and perceptual quality. Jimin Wang, Xianyang Wang, Deyang Liu |
CW | 5 |
| 2025 | PLGMNet: Parallel Local-Global Mamba Network for Real-Time Steel Surface Defect Detection
Chenlei Li, Xiaofei Zhou 0003, Yong Wu 0007, Deyang Liu, Jiyong Zhang 0001, Zhi Liu 0003 |
PRCV (17) | 5 |
| 2025 | Deep hierarchical network for full-reference omnidirectional image quality assessment
Youzhi Zhang 0004, Lifei Wan, Xiaofei Zhou 0003, Deyang Liu |
Multim. Syst. | 6 |
| 2025 | L3FMamba: Low-Light Light Field Image Enhancement With Prior-Injected State Space ModelsabstractIn this paper, we address the problem of low-light light field (LF) image enhancement, where spatial details and angular coherence are severely degraded due to noise and insufficient illumination. Existing methods often rely on local aggregation or naive view stacking, which fail to capture global illumination and long-range spatial-angular correlations. To overcome these limitations, we propose L3FMamba, a lightweight enhancement method that integrates Retinex and Atmospheric Scattering models with dark, bright, and average channel priors for robust illumination decomposition. Moreover, we incorporate a state space model to capture non-local spatial-angular dependencies, enabling effective propagation of global context across views. By combining physics-inspired priors with structured modeling, L3FMamba achieves accurate illumination correction and fine-detail preservation with minimal parameters. Experiments show that L3FMamba outperforms the state-of-the-art in quality. Deyang Liu, Shizheng Li, Zeyu Xiao 0002, Ping An 0001, Caifeng Shan |
IEEE Signal Process. Lett. | 1 |
| 2025 | Deep Sparse-to-Dense Inbetweening for Multi-View Light FieldsabstractLight field (LF) imaging, which captures both intensity and directional information of light rays, extends the capabilities of traditional imaging techniques. In this paper, we introduce a task in the field of LF imaging, sparse-to-dense inbetweening, which focuses on generating dense novel views from sparse multi-view LFs. By synthesizing intermediate views from sparse inputs, this task enhances LF view synthesis through filling in interperspective gaps within an expanded field of view and increasing data robustness by leveraging complementary information between light rays from different perspectives, which are limited by non-robust single-view synthesis and the inability to handle sparse inputs effectively. To address these challenges, we construct a high-quality multi-view LF dataset, consisting of 60 indoor scenes and 59 outdoor scenes. Building upon this dataset, we propose a baseline method. Specifically, we introduce an adaptive alignment module to dynamically align information by capturing relative displacements. Next, we explore angular consistency and hierarchical information using a multi-level feature decoupling module. Finally, a multi-level feature refinement module is applied to enhance features and facilitate reconstruction. Additionally, we introduce a universally applicable artifact-aware loss function to effectively suppress visual artifacts. Experimental results demonstrate that our method outperforms existing approaches, establishing a benchmark for sparse-to-dense inbetweening. The code is available at https://github.com/Starmao1/MutiLF. Zeyu Xiao 0002, Ping An 0001, Deyang Liu, Caifeng Shan |
IEEE Trans. Image Process. | 4 |
| 2024 | Low-Light Light-Field Image Enhancement With Geometry Consistency
Deyang Liu, Zhengqu Li, Jian Ma 0012 |
PRCV (8) | 1 |
| 2024 | Multiview Light Field Angular Super-Resolution Based on View Alignment and Frequency Attention
Deyang Liu, Youzhi Zhang 0004 |
PRCV (6) | 1 |
| 2024 | A foreground-context dual-guided network for light-field salient object detectionabstractLight-field salient object detection (SOD) has become an emerging trend as it records comprehensive information about natural scenes that can benefit salient object detection in various ways. However, salient object detection models with light-field data as input have not been thoroughly explored. The existing methods cannot effectively suppress the noise, and it is difficult to distinguish the foreground and background under challenging conditions including self-similarity, complex backgrounds, large depth of field, and non-Lambertian scenarios. In order to extract the feature of light-field images effectively and suppress the noise in light-field, in this paper, we propose a foreground and context dual guided network. Specifically, we design a global context extraction module (GCEM) and a local foreground extraction module (LFEM). GCEM is used to suppress global noise and roughly predict saliency maps. GCEM also can extract global context information from deep-level features to guide decoding process. By extracting local information from shallow-level, LFEM refines the prediction obtained by GCEM. In addition, we use RGB images to enhance the light-field images before the input GCEM. Experimental results show that our proposed method is effective in suppressing global noise and achieves better results when dealing with transparent objects and complex backgrounds. The experimental results show that the proposed method outperforms several other state-of-the-art methods on three light-field datasets. Xin Zheng 0006, Deyang Liu, Chengtao Lv, Jiebin Yan |
Signal Process. Image Commun. | 3 |
| 2024 | Spatial Attention-Guided Light Field Salient Object Detection Network With Implicit Neural RepresentationabstractRecently, many Light Field Salient Object Detection (LF SOD) methods have been proposed. However, guaranteeing the integrality and recovering more high-frequency details of the generated salient object map still remain challenging. To this end, we propose a spatial attention-guided LF SOD network with implicit neural representation to further improve LF SOD performance. We adopt an encoder-decoder structure for model construction. In order to ensure the completeness of the generated salient object map, a multi-modal and multi-scale feature fusion module is designed in the encoder part to refine the salient regions within all-in-focus image and aggregate the focal stack and all-in-focus image in spatial attention-guided manner. In order to recover more high-frequency details of the obtained salient object map, an implicit detail restoration module is proposed in the decoder part. In virtue of implicit neural representation, we convert the detail restoration problem into a functional mapping problem. By further integrating the self-attention mechanism, the derived saliency map can be depicted at a more refined level. Comprehensive experimental results demonstrate the superiority of the proposed method. Ablation studies and visual comparisons further validate that the proposed method can guarantee the integrality and recover more high-frequency detail information of the obtained saliency map. The code is publicly available athttps://github.com/ldyorchid/LFSOD-Net. Xin Zheng 0006, Zhengqu Li, Deyang Liu, Xiaofei Zhou 0003, Caifeng Shan |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2023 | Learning a Multilevel Cooperative View Reconstruction Network for Light Field Angular Super-ResolutionabstractRecently, many methods have been proposed to improve the angular resolution of sparsely-sampled Light Field (LF). However, the synthesized dense LF inevitably exhibits blurry edges and artifacts. This paper intents to model the global relations of LF views and quality degradation model by learning a multilevel cooperative view reconstruction network to further enhance LF angular Super-Resolution (SR) performance. The proposed LF angular SR network consists of three sub-networks including the Cooperative Angular Transformer Network (CATNet), the Deblurring Network (DBNet), and the Texture Repair Network (TRNet). The CATNet simultaneously captures global features of all LF views and local features within each view, which benefits in characterizing the inherent LF structure. The DBNet models a quality degradation model by estimating blur kernels to reduce the blurry edges and artifacts. The TRNet focuses on restoring fine-scale texture details. Experimental results over various LF datasets including large baseline LF images demonstrate the significant superiority of our method when compared with state-of-the-art ones. Deyang Liu, Xiaofei Zhou 0003, Ping An 0001, Yuming Fang 0001 |
ICME | 1 |
| 2023 | SRI-Net: Similarity retrieval-based inference network for light field salient object detection
Chengtao Lv, Xiaofei Zhou 0003, Deyang Liu, Bolun Zheng, Jiyong Zhang 0001, Chenggang Yan 0001 |
J. Vis. Commun. Image Represent. | 4 |
| 2023 | Geometry-assisted multi-representation view reconstruction network for Light Field image angular super-resolution
Deyang Liu, Zaidong Tong, Yan Huang 0023, Yifan Zuo 0001, Yuming Fang 0001 |
Knowl. Based Syst. | 1 |
| 2023 | Optical flow-assisted multi-level fusion network for Light Field image angular reconstruction
Deyang Liu, Yan Huang 0023, Yuanzhi Wang, Yuming Fang 0001 |
Signal Process. Image Commun. | 1 |
| 2023 | Gradient-Guided Single Image Super-Resolution Based on Joint Trilateral Feature FilteringabstractThe state of the arts (SOTAs) of single image super-resolution always exploit guidance from gradient prior. The fusion of gradient guidance is implemented by channel-wise concatenation followed by a convolutional layer. However, the kernels sharing in spatial positions cannot adaptively tune the effect of gradient guidance for all feature positions. To resolve this problem, a novel network module is proposed to simulate the traditional Joint Trilateral Filter (JTF) by extending the definition domain from pixels to features. Moreover, to improve the efficiency and flexibility, the functions of JTF kernel generation for image features and gradient features are explicitly learned instead of individual kernel weights, e.g., the exponential functions in the traditional JTF. Based on the proposed JTF modules, this paper follows the gradient-guided framework which simultaneously infers high-resolution (HR) image features and HR gradient features within two parallel branches, respectively. Specifically, by treating image features and gradient features as cross guidance to each other, the proposed JTF modules adaptively adjust the fusion patterns for local features via a bi-directional way. By doing so, the quality of image features and gradient features is alternatively enhanced. Compared with SOTAs, the proposed JTF-SISR shows improvement which is evaluated for multiple upsampling scales and degradation modes on 5 synthetic datasets, i.e., Set5, Set14, B100, Urban100 and Manga109, and 1 real dataset, i.e., RealSRSet. The code is public inhttps://github.com/a239xjc/JTF-SISR. Yifan Zuo 0001, Yuming Fang 0001, Deyang Liu, Wenying Wen |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2023 | Multi-Stream Dense View Reconstruction Network for Light Field Image CompressionabstractRecently, many view synthesis-based methods are proposed for high-efficiency light field (LF) image compression. However, most existing methods fail to recover more texture details on occlusion regions, which reduces the compression efficiency. In this paper, we propose a multi-stream dense view reconstruction network to further improve LF image compression performance. In our method, only sparsely-sampled LF views are transmitted and the rest of the views are reconstructed at the decoder side. During the reconstruction process, we firstly constitute a multi-disparity geometry (MDG) structure based on the decoded sparse LF views, which can reflect abundant disparity characteristics. Subsequently, a multi-stream view reconstruction network (MSVRNet) is put forward to reconstruct a high-quality dense LF image, which consists of a multi-scale feature fusion sub-network, a fusion reconstruction sub-network, and a detail refinement sub-network. The multi-scale feature fusion sub-network can implicitly lean abundant multiscale geometric structure features from the constituted MDG structure. The fusion reconstruction sub-network and the detail refinement sub-network are respectively utilized to fuse the learned multiscale geometric features and restore more texture details, especially for occlusion regions. Moreover, 3D convolutional operations are adopted in the whole reconstruction process, which allow information propagation among the learned multiscale geometric features. Comprehensive experimental results demonstrate the effectiveness of the proposed method. The perceptual quality of reconstructed views and application on depth estimation also demonstrate that the proposed method can keep structural consistency of the reconstructed LF image and recover more texture details. Deyang Liu, Yan Huang 0023, Yuming Fang 0001, Yifan Zuo 0001, Ping An 0001 |
IEEE Trans. Multim. | 1 |
| 2022 | On-Site Monitoring of Forging Press: An Application of Volapu IIoT SuiteabstractIn the process of implementing the transformation and upgrading of Industry 4.0, in order to protect the existing investment of manufacturing customers, we often need to do digitized transformation and upgrade of the existing equipment on the production site. In order to achieve this task quickly, we have designed and developed an out-of-the-box industrial Internet of Things system for industrial sites, Volapu IIoT Suite, which can be flexibly configured according to the needs of different production equipment. In this paper, we take the transformation of the Internet of Things of the forging press as an example, and give an on-site data collection plan for the forging press and an example of analysis of production data. Dashuai Guo, Po-Hsun Chueh, Deyang Liu, Joseph K. H. Wang |
SNPD | 3 |
| 2022 | Selective part-based correlation filter tracking algorithm with reinforcement learningabstractAbstract In visual object tracking methods, improving both the run time and the accuracy in the face of complex situations has always been an important issue. Many complex tracking algorithms, such as part‐based algorithms, have better accuracy when facing occlusions, but they have much greater computational complexity. In response to the above problems, this paper proposes a selective part‐based correlation filter (SPCF) tracking algorithm with a reinforcement learning to achieve more stable and efficient tracking of targets. First, according to the conditions of the response map of the correlation filter (CF), the entire tracking process is divided into three states: simple environments, complex environments, and harsh environments. Second, this paper uses reinforcement learning to determine the states of frames in different situations to improve the tracking effect of the algorithm. Third, the process of the online selection of states is transformed into a Markov decision process (MDP), where the policy learning of the MDP is achieved by reinforcement learning. Additionally, different strategies are used to track a target in different states: the overall filter is used to increase the speed in simple environments; part‐based filters are used to improve the accuracy in complex environments; and in harsh environments where the target completely disappears, a redetection algorithm is used to find the target when it reappears. Finally, the performance of the tracking algorithm is verified on the VOT2018, OTB‐2015, and LaSOT datasets. Zhengzhi Lu, Guoan Yang, Deyang Liu, Junjie Yang 0009, Chuanbo Zhou |
IET Image Process. | 3 |
| 2022 | Energy-driven reference selection for hierarchical light field compression
Xinpeng Huang, Ping An 0001, Deyang Liu |
Signal Process. Image Commun. | 4 |
| 2022 | Low Bitrate Light Field Compression With Geometry and Content ConsistencyabstractLight field imaging can simultaneously record the position and direction information of light rays; thus, digital refocusing and full depth-of-field extension — functions that are inaccessible for conventional images — can be achieved using the structural consistency of light field data. To meet the challenges of limited bandwidth and storage, such vast numbers of light field data must be compressed to a low bitrate. However, current compression solutions ignore the intrinsic consistency of light fields in pursuit of a low bitrate, thereby leading to the loss of light field capabilities. To solve this issue, this work focuses on structural consistency to achieve efficient light field compression with a low bitrate. The proposed light field compression method encodes the sparsely selected sub-aperture images (SAIs) and the disparity maps corresponding to the unselected SAIs. From the perspective of geometry consistency, the consistency of the initially estimated disparity maps is improved by using a color-guided refinement algorithm, thereby reducing the bitrate of the disparity maps. From the perspective of content consistency, the consistency of the SAI-transformed pseudo sequence is improved by the proposed content-similarity-based arrangement algorithm along with a specific prediction structure; thereby, the bitrate of the sparsely selected SAIs is reduced. The experimental results show that the proposed compression method can reduce the total bitrate while preserving good structural consistency. Xinpeng Huang, Ping An 0001, Deyang Liu, Liquan Shen |
IEEE Trans. Multim. | 4 |
| 2021 | View synthesis-based light field image compression using a generative adversarial network
Deyang Liu, Xinpeng Huang, Wenfa Zhan, Liefu Ai, Shulin Cheng |
Inf. Sci. | 1 |
| 2021 | Learning from EPI-Volume-Stack for Light Field image angular super-resolution
Deyang Liu, Qiang Wu 0001, Yan Huang 0023, Xinpeng Huang, Ping An 0001 |
Signal Process. Image Commun. | 1 |
| 2020 | A convolutional neural network with sparse representation
Guoan Yang, Zhengzhi Lu, Deyang Liu |
Knowl. Based Syst. | 4 |
| 2020 | Light Field Compression Using Global Multiplane Representation and Two-Step PredictionabstractDue to its spatio-angular structure, light field image allows for a wealth of post-processing techniques like digital refocusing and depth estimation. In order to compress the data of the two domains, the current proposal intends to embed the disparity-based view synthesis method into the decoder. However, predicting each view separately or in local groups means bringing more computational burden to the decoder and destroying the light field structure. Since disparity contains the relationship between all light rays in the light field, the proposed solution is to predict a disparity-based global representation as the first step. In the second step, all the views can be predicted easily based on this representation. In this letter, we use the recently proposed multiplane as the form of this global representation. The experimental results show the effectiveness of the proposed solution, and the better RD performance compared to other schemes especially under low bitrates. Ping An 0001, Xinpeng Huang, Chao Yang 0021, Deyang Liu, Qiang Wu 0001 |
IEEE Signal Process. Lett. | 5 |
| 2020 | Full Reference Light Field Image Quality Evaluation Based on Angular-Spatial CharacteristicabstractThe quality evaluation is an indispensable link in light field (LF) image processing. Most of existing LF objective evaluation indexes do not make effective use of the angular characteristic of LF, so the evaluation results are unsatisfactory. In this letter, the quality evaluation of LF image is constructed based on human visual system (HVS) and LF angular-spatial characteristics. Based on the fact that HVS has different sensitivity to different parallaxes, we assume that LF image quality perceived by human eyes has the optimal parallax range. A dual-fan filter is used to constrain the parallax range. Then, the overall quality of the LF is represented by combining the spatial and angular quality, which performed from the central sub-aperture image and the focus stack, respectively. In addition, because the difference of Gaussian (DoG) operator can simulate the process of extracting texture structure by human eyes. The structural similarity of DoG texture feature is utilized in the spatial quality evaluation. Extensive comparison experiments show that the proposed method is more consistent with the characteristics of LF. Chunli Meng, Ping An 0001, Xinpeng Huang, Chao Yang 0021, Deyang Liu |
IEEE Signal Process. Lett. | 5 |
| 2020 | Content-Based Light Field Image Compression Method With Gaussian Process RegressionabstractLight field (LF) imaging enables new possibilities for digital imaging, such as digital refocusing, changing of focus plane, changing of viewpoint, scene-depth estimation, and 3D scene reconstruction, by capturing both spatial and angular information of light rays. However, one main problem in dealing with LF data is its sheer volume. In this context, efficient compression methods are needed for such a particular type of content. In this paper, we propose a content-based LF image-compression method with Gaussian process regression to improve the compression efficiency and accelerate the prediction procedure. First, the LF image is fed to the intra-frame codec of HEVC. In the prediction procedure, the prediction units (PUs) are classified as non-homogenous texture units, homogenous texture units, and visually flat units, based on the content property of the LF image. For each category, we design a corresponding Gaussian process regression (GPR)-based prediction method. Moreover, we propose a classification mechanism to exactly decide to which category the current PU belongs, so as to adjust the trade-off between the computational burden and the LF image coding efficiency. Experimental results demonstrate that the proposed LF image compression method is superior to several other state-of-the-art compression methods in terms of different quality metrics. Furthermore, the proposed method can also achieve a good visual quality of views rendered from decoded LF contents. Deyang Liu, Ping An 0001, Wenfa Zhan, Xinpeng Huang, Ali Abdullah Yahya |
IEEE Trans. Multim. | 1 |
| 2018 | Hybrid linear weighted prediction and intra block copy based light field image coding
Deyang Liu, Ping An 0001, Liquan Shen |
Multim. Tools Appl. | 1 |
| 2018 | Scalable coding of 3D holoscopic image by using a sparse interlaced view image set and disparity map
Deyang Liu, Ping An 0001, Chao Yang 0021, Liquan Shen, Kai Li 0016 |
Multim. Tools Appl. | 1 |
| 2017 | Coding of 3D holoscopic image by using spatial correlation of rendered view imagesabstractHoloscopic imaging is a prospective acquisition and display solution for providing natural and fatigue-free 3D visualization. However, large amount of data is required to represent the 3D holoscopic content. Therefore, efficient coding schemes for this particular type of image are needed. In this paper, an effective coding scheme is proposed by exploring the spatial correlation among the view images with different perspectives rendered from 3D holoscopic image. We utilize the interlaced view image to descript such spatial correlation. A linear prediction method is used on the interlaced view image instead of the original holoscopic image directly. Experimental results show that the proposed coding scheme performs better than HEVC intra standard and screen content coding extension of HEVC with around 2.41dB and 0.42 dB average quality improvement respectively. Deyang Liu, Ping An 0001, Chao Yang 0021, Liquan Shen |
ICASSP | 1 |
| 2017 | Bit allocation for 3D video coding based on lagrangian multiplier adjustment
Chao Yang 0021, Ping An 0001, Deyang Liu, Liquan Shen, Kai Li 0016 |
Signal Process. Image Commun. | 3 |
| 2016 | Depth map coding based on virtual view qualityabstractMulti-view video plus depth (MVD) is a 3D video representation. In MVD, the depth map provides the scene distance information and is used to render the virtual view through Depth Image Based Rendering (DIBR) technique. The depth map coding error will induce distortion in the rendered virtual views. This paper proposes a mathematic model that can estimate the synthesized virtual view distortion induced by depth map compression, and the model is employed to the rate distortion optimization (RDO) in the depth map coding. Based on the rendered virtual view quality, a Lagrangian optimization adjustment scheme at Coding Unit (CU) level is proposed to improve the depth map encoding efficiency. Experimental results demonstrate that the proposed method can improve the BD-PSNR of virtual view for 0.62 dB, and the encoding complexity reduces compared with the view synthesis optimization (VSO) technique in the 3D-HEVC Test Model (HTM). Chao Yang 0021, Ping An 0001, Deyang Liu, Liquan Shen |
ICASSP | 3 |
| 2016 | 3D holoscopic image coding scheme using HEVC with Gaussian process regression
Deyang Liu, Ping An 0001, Chao Yang 0021, Liquan Shen |
Signal Process. Image Commun. | 1 |
| 2015 | Virtual view distortion estimation for depth map codingabstractMulti-view video plus depth (MVD) format is a three-dimensional (3D) video representation. The depth map in MVD provides the scene geometry information and is used to render the virtual view through Depth Image Based Rendering (DIBR). In this paper, a virtual view distortion estimation function based on the characteristics of both texture image and depth map is proposed which can estimate virtual view distortion induced by depth map compression accurately, and the function is implemented to the Rate Distortion Optimization (RDO) in the depth map coding. Compared with the View Synthesis Optimization (VSO) in 3D-HEVC Test Model (HTM) reference software, the experimental results demonstrate that the proposed method can improve the BD-PSNR of virtual view for 0.26 dB on average, and the encoding time has reduced for 31% on average due to the low complexity of the proposed function. Chao Yang 0021, Ping An 0001, Deyang Liu, Liquan Shen |
VCIP | 3 |