Shansi Zhang

dblp:266/0158 · DBLP profile ↗
← Back
8ranked-venue papers
8as first author
8since 2021 · last 2024
0000-0002-2441-7615ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 7 first-author · 7 since 2021Artificial intelligence and machine learning · 2 · 2 first-author · 2 since 2021
YearPublicationVenuePosition
2024 Unsupervised Disparity Estimation for Light Field Videos
abstract
Light field (LF) videos contain not only the spatial-angular information but also the temporal information, which are useful for disparity estimation. The existing work on disparity estimation for LF videos relies on supervised training with disparity labels. To overcome this reliance, we develop an unsupervised disparity estimation framework for LF videos, which consists of a matching branch to perform feature matching and a refinement branch to refine the disparity maps. Our framework also includes a cross-feature fusion module with self-attention and cross-attention to fuse the multi-frame features, and a cost aggregation transformer with cross-depth self-attention blocks to explore the global depth dependencies. Moreover, we propose a left-right consistency strategy to estimate the occlusion regions for the input views and introduce a occlusion-aware photometric loss to solve the occlusion issue. Experimental results demonstrate that our method achieves superior performance compared to the existing supervised and unsupervised methods.
Shansi Zhang, Edmund Y. Lam
ICASSP1
2024 Light Field Image Restoration via Latent Diffusion and Multi-View Attention
abstract
Light field (LF) images contain information for multiple views. The restoration of degraded LF images is of great significance for various LF applications. Inspired by the recent achievement of denoising diffusion models, we propose a LF image restoration method based on latent diffusion (LD). We design a LDUNet with efficient cross-attention modules to integrate the features of conditional input, and propose a two-stage training strategy, where the LDUNet is first trained on the individual views and then fine-tuned on the LF images with injected prior noise. A refinement module is jointly trained in the second stage to enhance the spatial-angular structures. It consists of multi-view attention blocks with patch-based angular self-attention to fuse the global view information. Moreover, we introduce an enhanced noise loss for better noise prediction and an auxiliary image loss to obtain high-quality images. We evaluate our method on LF image deraining task and low-light LF image enhancement task. Our method demonstrates superior performance on both tasks compared to the existing methods.
Shansi Zhang, Edmund Y. Lam
IEEE Signal Process. Lett.1
2024 Unsupervised Light Field Depth Estimation via Multi-View Feature Matching With Occlusion Prediction
abstract
Depth estimation from light field (LF) images is a fundamental step for numerous applications. Recently, learning-based methods have achieved higher accuracy and efficiency than the traditional methods. However, it is costly to obtain sufficient depth labels for supervised training. In this paper, we propose an unsupervised framework to estimate depth from LF images. First, we design a disparity estimation network (DispNet) with a coarse-to-fine structure to predict disparity maps from different view combinations. It explicitly performs multi-view feature matching to learn the correspondences effectively. As occlusions may cause the violation of photo-consistency, we introduce an occlusion prediction network (OccNet) to predict the occlusion maps, which are used as the element-wise weights of photometric loss to solve the occlusion issue and assist the disparity learning. With the disparity maps estimated by multiple input combinations, we then propose a disparity fusion strategy based on the estimated errors with effective occlusion handling to obtain the final disparity map with higher accuracy. Experimental results demonstrate that our method achieves superior performance on both the dense and sparse LF images, and also shows better robustness and generalization on the real-world LF images compared to the other methods.
Shansi Zhang, Nan Meng, Edmund Y. Lam
IEEE Trans. Circuits Syst. Video Technol.1
2024 Semi-Supervised Semantic Segmentation for Light Field Images Using Disparity Information
abstract
Light field (LF) images enable numerous applications due to their ability to capture information for multiple views. Semantic segmentation is an essential task for LF scene understanding. However, existing supervised methods heavily rely on a large number of pixel-wise annotations. To relieve this problem, we propose a semi-supervised LF semantic segmentation method that requires only a small subset of labeled data and harnesses the LF disparity information. First, we design an unsupervised disparity estimation network, which can determine the disparity map for every view. With the estimated disparity maps, we generate pseudo-labels along with their weight maps for the peripheral views when only the labels of central views are available. We then merge the predictions from multiple views to obtain more reliable pseudo-labels for unlabeled data, and introduce a disparity-semantics consistency loss to enforce structure similarity. Moreover, we develop a comprehensive contrastive learning scheme that includes a pixel-level strategy to enhance feature representations and an object-level strategy to improve segmentation for individual objects. Our method demonstrates state-of-the-art performance on the benchmark LF semantic segmentation dataset under a variety of training settings and achieves comparable performance to supervised methods when trained under 1/2 protocol.
Shansi Zhang, Edmund Y. Lam
IEEE Trans. Image Process.1
2023 LRT: An Efficient Low-Light Restoration Transformer for Dark Light Field Images
abstract
Light field (LF) images containing information for multiple views have numerous applications, which can be severely affected by low-light imaging. Recent learning-based methods for low-light enhancement have some disadvantages, such as a lack of noise suppression, complex training process and poor performance in extremely low-light conditions. To tackle these deficiencies while fully utilizing the multi-view information, we propose an efficient Low-light Restoration Transformer (LRT) for LF images, with multiple heads to perform intermediate tasks within a single network, including denoising, luminance adjustment, refinement and detail enhancement, achieving progressive restoration from small scale to full scale. Moreover, we design an angular transformer block with an efficient view-token scheme to model the global angular dependencies, and a multi-scale spatial transformer block to encode the multi-scale local and global information within each view. To address the issue of insufficient training data, we formulate a synthesis pipeline by simulating the major noise sources with the estimated noise parameters of LF camera. Experimental results demonstrate that our method achieves the state-of-the-art performance on low-light LF restoration with high efficiency.
Shansi Zhang, Nan Meng, Edmund Y. Lam
IEEE Trans. Image Process.1
2022 A Deep Retinex Framework for Light Field Restoration under Low-light Conditions
abstract
Light field (LF) images can record the scene from multiple directions and have many applications, such as refocusing and depth estimation. However, these applications can be heavily influenced by poor light condition and noise. This work aims to recover the high-quality LF images from their lowlight detection. First, a decomposition network is employed to decompose each LF image into its reflectance and illumination with the Retinex theory. Then, two enhancement networks are designed to denoise the reflectance and enhance the illumination, respectively. They adopt alternate spatial-angular feature extractions and process all the views synchronously with high efficiency. A parallel dual attention mechanism is integrated to both the spatial and angular feature extractions, to encode more important information. Moreover, a discriminator is introduced during the training to generate more realistic LF images by making judgment according to both the spatial and angular characteristics. Experimental results have demonstrated the superior performance of our method, which can restore the content, luminance, color and geometric structures of LF images effectively.
Shansi Zhang, Edmund Y. Lam
ICPR1
2021 Learning to restore light fields under low-light imaging
Shansi Zhang, Edmund Y. Lam
Neurocomputing1
2021 An effective decomposition-enhancement method to restore light field images captured in the dark
Shansi Zhang, Edmund Y. Lam
Signal Process.1