VLDB 2026 Research / reviewers in the wild / expert
Shansi Zhang
dblp:266/0158
· DBLP profile ↗
8ranked-venue papers
8as first author
8since 2021 · last 2024
0000-0002-2441-7615ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 7 first-author · 7 since 2021Artificial intelligence and machine learning · 2 · 2 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Unsupervised Disparity Estimation for Light Field VideosabstractLight field (LF) videos contain not only the spatial-angular information but also the temporal information, which are useful for disparity estimation. The existing work on disparity estimation for LF videos relies on supervised training with disparity labels. To overcome this reliance, we develop an unsupervised disparity estimation framework for LF videos, which consists of a matching branch to perform feature matching and a refinement branch to refine the disparity maps. Our framework also includes a cross-feature fusion module with self-attention and cross-attention to fuse the multi-frame features, and a cost aggregation transformer with cross-depth self-attention blocks to explore the global depth dependencies. Moreover, we propose a left-right consistency strategy to estimate the occlusion regions for the input views and introduce a occlusion-aware photometric loss to solve the occlusion issue. Experimental results demonstrate that our method achieves superior performance compared to the existing supervised and unsupervised methods. Shansi Zhang, Edmund Y. Lam |
ICASSP | 1 |
| 2024 | Light Field Image Restoration via Latent Diffusion and Multi-View AttentionabstractLight field (LF) images contain information for multiple views. The restoration of degraded LF images is of great significance for various LF applications. Inspired by the recent achievement of denoising diffusion models, we propose a LF image restoration method based on latent diffusion (LD). We design a LDUNet with efficient cross-attention modules to integrate the features of conditional input, and propose a two-stage training strategy, where the LDUNet is first trained on the individual views and then fine-tuned on the LF images with injected prior noise. A refinement module is jointly trained in the second stage to enhance the spatial-angular structures. It consists of multi-view attention blocks with patch-based angular self-attention to fuse the global view information. Moreover, we introduce an enhanced noise loss for better noise prediction and an auxiliary image loss to obtain high-quality images. We evaluate our method on LF image deraining task and low-light LF image enhancement task. Our method demonstrates superior performance on both tasks compared to the existing methods. Shansi Zhang, Edmund Y. Lam |
IEEE Signal Process. Lett. | 1 |
| 2024 | Unsupervised Light Field Depth Estimation via Multi-View Feature Matching With Occlusion PredictionabstractDepth estimation from light field (LF) images is a fundamental step for numerous applications. Recently, learning-based methods have achieved higher accuracy and efficiency than the traditional methods. However, it is costly to obtain sufficient depth labels for supervised training. In this paper, we propose an unsupervised framework to estimate depth from LF images. First, we design a disparity estimation network (DispNet) with a coarse-to-fine structure to predict disparity maps from different view combinations. It explicitly performs multi-view feature matching to learn the correspondences effectively. As occlusions may cause the violation of photo-consistency, we introduce an occlusion prediction network (OccNet) to predict the occlusion maps, which are used as the element-wise weights of photometric loss to solve the occlusion issue and assist the disparity learning. With the disparity maps estimated by multiple input combinations, we then propose a disparity fusion strategy based on the estimated errors with effective occlusion handling to obtain the final disparity map with higher accuracy. Experimental results demonstrate that our method achieves superior performance on both the dense and sparse LF images, and also shows better robustness and generalization on the real-world LF images compared to the other methods. Shansi Zhang, Nan Meng, Edmund Y. Lam |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2024 | Semi-Supervised Semantic Segmentation for Light Field Images Using Disparity InformationabstractLight field (LF) images enable numerous applications due to their ability to capture information for multiple views. Semantic segmentation is an essential task for LF scene understanding. However, existing supervised methods heavily rely on a large number of pixel-wise annotations. To relieve this problem, we propose a semi-supervised LF semantic segmentation method that requires only a small subset of labeled data and harnesses the LF disparity information. First, we design an unsupervised disparity estimation network, which can determine the disparity map for every view. With the estimated disparity maps, we generate pseudo-labels along with their weight maps for the peripheral views when only the labels of central views are available. We then merge the predictions from multiple views to obtain more reliable pseudo-labels for unlabeled data, and introduce a disparity-semantics consistency loss to enforce structure similarity. Moreover, we develop a comprehensive contrastive learning scheme that includes a pixel-level strategy to enhance feature representations and an object-level strategy to improve segmentation for individual objects. Our method demonstrates state-of-the-art performance on the benchmark LF semantic segmentation dataset under a variety of training settings and achieves comparable performance to supervised methods when trained under 1/2 protocol. Shansi Zhang, Edmund Y. Lam |
IEEE Trans. Image Process. | 1 |
| 2023 | LRT: An Efficient Low-Light Restoration Transformer for Dark Light Field ImagesabstractLight field (LF) images containing information for multiple views have numerous applications, which can be severely affected by low-light imaging. Recent learning-based methods for low-light enhancement have some disadvantages, such as a lack of noise suppression, complex training process and poor performance in extremely low-light conditions. To tackle these deficiencies while fully utilizing the multi-view information, we propose an efficient Low-light Restoration Transformer (LRT) for LF images, with multiple heads to perform intermediate tasks within a single network, including denoising, luminance adjustment, refinement and detail enhancement, achieving progressive restoration from small scale to full scale. Moreover, we design an angular transformer block with an efficient view-token scheme to model the global angular dependencies, and a multi-scale spatial transformer block to encode the multi-scale local and global information within each view. To address the issue of insufficient training data, we formulate a synthesis pipeline by simulating the major noise sources with the estimated noise parameters of LF camera. Experimental results demonstrate that our method achieves the state-of-the-art performance on low-light LF restoration with high efficiency. Shansi Zhang, Nan Meng, Edmund Y. Lam |
IEEE Trans. Image Process. | 1 |
| 2022 | A Deep Retinex Framework for Light Field Restoration under Low-light ConditionsabstractLight field (LF) images can record the scene from multiple directions and have many applications, such as refocusing and depth estimation. However, these applications can be heavily influenced by poor light condition and noise. This work aims to recover the high-quality LF images from their lowlight detection. First, a decomposition network is employed to decompose each LF image into its reflectance and illumination with the Retinex theory. Then, two enhancement networks are designed to denoise the reflectance and enhance the illumination, respectively. They adopt alternate spatial-angular feature extractions and process all the views synchronously with high efficiency. A parallel dual attention mechanism is integrated to both the spatial and angular feature extractions, to encode more important information. Moreover, a discriminator is introduced during the training to generate more realistic LF images by making judgment according to both the spatial and angular characteristics. Experimental results have demonstrated the superior performance of our method, which can restore the content, luminance, color and geometric structures of LF images effectively. Shansi Zhang, Edmund Y. Lam |
ICPR | 1 |
| 2021 | Learning to restore light fields under low-light imaging
Shansi Zhang, Edmund Y. Lam |
Neurocomputing | 1 |
| 2021 | An effective decomposition-enhancement method to restore light field images captured in the dark
Shansi Zhang, Edmund Y. Lam |
Signal Process. | 1 |