EDBT 2026 Demo / reviewers in the wild / expert
Shuo Zhang 0003
dblp:83/3714-3
· DBLP profile ↗
38ranked-venue papers
8as first author
26since 2021 · last 2026
0000-0003-4622-0669ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 27 · 7 first-author · 19 since 2021Artificial intelligence and machine learning · 16 · 4 first-author · 11 since 2021Databases, data management, data science and information retrieval · 2Computer networks · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Efficient multi-agent communication via entity-aware causal network
Yifan Bo, Jinghan Feng, Shuo Zhang 0003, Biao Leng |
Neural Networks | 4 |
| 2026 | Consistency guided multiple plane image construction for novel view synthesis
Yichang Lv, Shuo Zhang 0003, Youfang Lin |
Neural Networks | 2 |
| 2026 | Scattering center guided mono-static radar cross section prediction
Zehao Tang, Shuo Zhang 0003, Biao Leng |
Neural Networks | 3 |
| 2026 | A unified occlusion-free framework for unsupervised light field depth estimation
Longzhao Guo, Shuo Zhang 0003, Youfang Lin |
Pattern Recognit. | 2 |
| 2026 | DuoNet: Joint optimization of representation learning and prototype classifier for unbiased scene graph generation
Zhaodi Wang 0003, Biao Leng, Shuo Zhang 0003 |
Pattern Recognit. | 3 |
| 2026 | Focus-then-fusion: Learning discriminative cross-modal prototypes for few-shot classification
Rongshan Chen, Shuo Zhang 0003, Biao Leng |
Pattern Recognit. | 3 |
| 2026 | Hierarchical Interactive Multi-Plane Image Construction for Light Field Background and Reflection SeparationabstractReflections pose a significant challenge to the quality of light field images. Existing reflection removal methods typically regard it as pixel classification of 2D images, overlooking the fundamental principle that reflections arise from the 3D spatial superposition of background and reflection spaces. In this work, we propose a hierarchical interactive Multi-Plane Image (MPI) construction approach, which separates the mixed 3D space containing reflections and constructs two independent MPIs. We leverage the layered structure of MPIs to introduce three hierarchical interaction mechanisms: The inter interaction is designed to separate and recover background and reflection components, the intra interaction aims to reduce errors in information distribution within each plane, and the inner interaction focuses on optimizing the MPI structure itself. Ultimately, we successfully separate the mixed images and reconstruct two independent spatial representations. Compared to existing reflection removal methods, our approach not only achieves superior separation performance but also supports novel view synthesis of the separated results. In challenging scenarios with severe overlap between background and reflection, our method demonstrates a remarkable ability to improve separation quality. Shuo Zhang 0003, Yichang Lv, Youfang Lin |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2026 | Decoupling and Aggregating: Dual-Layer Light Field Depth Estimation With Reflective and Transparent SurfacesabstractLight Field (LF) is extensively utilized for depth estimation tasks due to its rich structural information. However, real-world LF images often encounter reflective and transparent surfaces and the related regions contain depth information from the reflection and background layers, which can be modeled as dual-layer scenes. For the existing depth estimation frameworks, the constructed cost volume shows an aliasing bimodal distribution in dual-layer surfaces and further causes serious wrong depth results. In this paper, we propose a novel decoupling-and-aggregating strategy and develop a dual-layer depth estimation network for LF images with complex reflections. Specifically, we develop an adaptive cost volume decoupling module to separate both the background and reflection features from the aliasing cost volume. Light field angular-spatial information is sufficiently extracted to infer the effort of features in different dimensions to the background or reflection layer. Additionally, we employ an iterative self-guided aggregating module with multi-stage supervision to aggregate two branches of cost volumes. The module applies the self-guided masks to regularize the distribution of cost volumes. Given the challenge of acquiring the ground truth disparity maps for the LF images under reflection scenes, we also construct a synthetic dataset with dual-layer properties. Our model is the first to introduce dual-layer scenes into the LF depth estimation task using an end-to-end deep neural network. It successfully separates the background and reflection layers and achieves accurate depth estimation results in both layers. Quantitative and qualitative experiment results on publicly available datasets demonstrate that our method performs better than other state-of-the-art methods. Shuo Zhang 0003, Yanlin Xie, Youfang Lin |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2026 | SPANext: Subpattern-Aware Two-Stage Graph Learning Framework for Next Location Prediction
Meiyue You, Shuo Zhang 0003, Biao Leng |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2026 | CrossHypergraph: Consistent High-Order Semantic Network for Few-Shot Image ClassificationabstractFew-shot classification is a challenging task that recognizes novel classes by learning from few training instances. Metric-based models are currently the most effective solutions for few-shot classification. In these models, patch feature distances between query instances and support classes are calculated to achieve classification. However, it is difficult for patch-based methods to mine semantic information of support and query instances, leading to inaccurate feature similarity measures. To address these problems, we propose to construct CrossHypergraph based on hypergraph modeling. Specifically, we first align the local prototype vertices of support and query instances to model consistent hypergraph structures. Then avertex-hyperedge-vertex-based interactive feature updating mechanism is designed to generate CrossHypergraph representation with consistent high-order semantic information for support and query instances. Based on the CrossHypergraph, we propose a consistent high-order semantic network, in which the high-order semantic-based weighted metric strategy is designed to achieve accurate classification. The proposed method is evaluated on general, fine-grained, and cross-domain few-shot benchmarks, including miniImageNet, tieredImageNet, CIFAR-FS, FC100, and miniImageNet$\rightarrow$CUB datasets. Experimental results show that our CrossHypergraph-based few-shot classifier generates consistent high-order semantic features, and achieves state-of-the-art performance on both 1-shot and 5-shot tasks. Shuo Zhang 0003, Biao Leng |
IEEE Trans. Multim. | 3 |
| 2025 | Unlocking the Potential of Reverse Distillation for Anomaly DetectionabstractKnowledge Distillation (KD) is a promising approach for unsupervised Anomaly Detection (AD). However, the student network's over-generalization often diminishes the crucial representation differences between teacher and student in anomalous regions, leading to detection failures. To address this problem, the widely accepted Reverse Distillation (RD) paradigm designs the asymmetry teacher and student network, using an encoder as teacher and a decoder as student. Yet, the design of RD does not ensure that the teacher encoder effectively distinguishes between normal and abnormal features or that the student decoder generates anomaly-free features. Additionally, the absence of skip connections results in a loss of fine details during feature reconstruction. To address these issues, we propose RD with Expert, which introduces a novel Expert-Teacher-Student network for simultaneous distillation of both the teacher encoder and student decoder. The added expert network enhances the student's ability to generate normal features and optimizes the teacher's differentiation between normal and abnormal features, reducing missed detections. Additionally, Guided Information Injection is designed to filter and transfer features from teacher to student, improving detail reconstruction and minimizing false positives. Experiments on several benchmarks prove that our method outperforms existing unsupervised AD methods under RD paradigm, fully unlocking RD’s potential. Biao Leng, Shuo Zhang 0003 |
AAAI | 4 |
| 2025 | Epipolar Consistent Attention Aggregation Network for Unsupervised Light Field Disparity Estimation
Shuo Zhang 0003, Youfang Lin |
ICCV | 2 |
| 2025 | Exploring View Consistency for Scene-Adaptive Low-Light Light Field Image Enhancement
Shuo Zhang 0003, Youfang Lin |
ICCV | 1 |
| 2025 | Epipolar Consistency-based Network for Structure-Aware LF Semantic SegmentationabstractLight Field (LF) semantic segmentation relies on leveraging redundant information across multiple views to assign a semantic label to each pixel of the central view. Recent approaches typically feed the views into a pre-trained backbone and utilize an estimated depth map to aggregate semantic representations for label prediction. However, these methods ignore the correlation between encoded structural cues in LF and semantic labels. On one hand, it is challenging to identify matching points for regions that are occluded in some views. This broken view consistency emphasizes object edge localization, facilitating more precise edge labeling. On the other hand, the depth continuity for the same object ensures semantic consistency in adjacent regions. Therefore, effectively extracting structural cues and integrating them into semantic segmentation are key points in LF semantic segmentation.In this paper, we propose an Epipolar Consistency-based network for structure-aware LF semantic segmentation, termed ECNet. First, we explore the epipolar consistency between views to characterize the edges and depth cues of the input. Based on the embedded edges information, we design an edge-semantic correlation transformer to generate fine-grained representations of object edges. Furthermore, the proposed depth-semantic correlation transformer maps semantic features of one object closer together according the depth information.Extensive experiments demonstrate that ECNet achieves state-of-the-art performance, which reduces computational cost by 33.3% (in terms of FLOPs) while maintaining high segmentation accuracy. Youfang Lin, Shuo Zhang 0003 |
ACM Multimedia | 4 |
| 2025 | Structure-Aware Pre-Selected Neural Rendering for Light Field ReconstructionabstractAs densely-sampled Light Field (LF) images are beneficial to many applications, LF reconstruction becomes an important technology in related fields. Recently, neural rendering shows great potential in reconstruction tasks. However, volume rendering in existing methods needs to sample many points on the whole camera ray or epipolar line, which is time-consuming. In this paper, specifically for LF images with regular angular sampling, we propose a novel Structure-Aware Pre-Selected neural rendering framework for LF reconstruction. Instead of sampling on the whole epipolar line, we propose to sample on several specific positions, which are estimated using the color and inherent scene structure information explored in the regular angular sampled LF images. Sampling only a few points that closely match the target pixel, the feature of the target pixel is quickly rendered with high-quality. Finally, we fuse the features and decode them in the view dimension to obtain the final target view. Experiments show that the proposed method outperforms the state-of-the-art LF reconstruction methods in both qualitative and quantitative comparisons across various tasks. Our method also surpasses the most existing methods in terms of speed. Moreover, without any retraining or fine-tuning, the performance of our method with no-per-scene optimization is even better than the methods with per-scene optimization. Song Chang, Youfang Lin, Shuo Zhang 0003 |
IEEE Trans. Multim. | 3 |
| 2025 | HGFormer: Topology-Aware Vision Transformer With HyperGraph LearningabstractThe computer vision community has witnessed an extensive exploration of vision transformers in the past two years. Drawing inspiration from traditional schemes, numerous works focus on introducing vision-specific inductive biases. However, the implicit modeling of permutation invariance and fully-connected interaction with individual tokens disrupts the regional context and spatial topology, further hindering higher-order modeling. This deviates from the principle of perceptual organization that emphasizes the local groups and overall topology of visual elements. Thus, we introduce the concept of hypergraph for perceptual exploration. Specifically, we propose a topology-aware vision transformer called HyperGraph Transformer (HGFormer). Firstly, we present a Center Sampling K-Nearest Neighbors (CS-KNN) algorithm for semantic guidance during hypergraph construction. Secondly, we present a topology-aware HyperGraph Attention (HGA) mechanism that integrates hypergraph topology as perceptual indications to guide the aggregation of global and unbiased information during hypergraph messaging. Using HGFormer as visual backbone, we develop an effective and unitive representation, achieving distinct and detailed scene depictions. Empirical experiments show that the proposed HGFormer achieves competitive performance compared to the recent SoTA counterparts on various visual benchmarks. Extensive ablation and visualization studies provide comprehensive explanations of our ideas and contributions. Shuo Zhang 0003, Biao Leng |
IEEE Trans. Multim. | 2 |
| 2025 | Progressive Multi-Plane Images Construction for Light Field Occlusion RemovalabstractRecently, Light Field (LF) shows great potential in removing occlusion since the objects occluded in some views may be visible in other views. However, existing LF-based methods implicitly model each scene and can only remove objects that have positive disparities in one central views. In this article, we propose a novel Progressive Multi-Plane Images (MPI) Construction method specifically designed for LF-based occlusion removal. Different from the previous MPI construction methods, we progressively construct MPIs layer by layer in order from near to far. In order to accurately model the current layer, the positions of foreground occlusions in the nearer layers are taken as occlusion prior. Specifically, we propose an Occlusion-Aware Attention Network to generate each layer of MPIs with reliable information in occluded regions. For each layer, occlusions in the current layer are filtered out so that the background is better recovered just using the visible views instead of the other occluded views. Then, by simply removing the layers containing occlusions and rendering MPIs in kinds of viewpoints, the occlusion removal results for different views are generated. Experiments on synthetic and real-world scenes show that our method outperforms state-of-the-art LF occlusion removal methods in quantitative and visual comparisons. Moreover, we also apply the proposed progressive MPI construction method to the view synthesis task. The occlusion edges in our synthesized views achieve significantly better quality, which also verifies that our method can better model the occluded regions. Shuo Zhang 0003, Song Chang, Zhuoyu Shi, Youfang Lin |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2024 | Dual-Modeling Decouple Distillation for Unsupervised Anomaly DetectionabstractKnowledge distillation based on student-teacher network is one of the mainstream solution paradigms for the challenging unsupervised Anomaly Detection task, utilizing the difference in representation capabilities of the teacher and student networks to implement anomaly localization. However, over-generalization of the student network to the teacher network may lead to negligible differences in representation capabilities of anomaly, thus affecting the detection effectiveness. Existing methods address the possible over-generalization by using differentiated students and teachers from the structural perspective or explicitly expanding distilled information from the content perspective, which inevitably results in an increased likelihood of underfitting of the student network and poor anomaly detection capabilities in anomaly center or edge. In this paper, we propose Dual-Modeling Decouple Distillation (DMDD) for the unsupervised Anomaly Detection. In DMDD, a Decouple Student-Teacher Network is proposed to decouple the initial student features into normality and abnormality features. We further introduce Dual-Modeling Distillation based on normal-anomalous image pairs, fitting normality features of anomalous image and the teacher features of the corresponding normal image, widening the distance between abnormality features and the teacher features in anomalous regions. Synthesizing these two distillation ideas, we achieve anomaly detection which focuses on both edge and center of anomaly. Finally, a Multi-perception Segmentation Network is proposed to achieve focused anomaly map fusion based on multiple attention. Experimental results on MVTec AD show that DMDD surpasses SOTA localization performance of previous knowledge distillation-based methods, reaching 98.85% on pixel-level AUC and 96.13% on PRO. Biao Leng, Shuo Zhang 0003 |
ACM Multimedia | 4 |
| 2024 | Multi-3D Occlusion Mask Learning for Flexible Occlusion Removal in Neural Radiance Fields
Zhuoyu Shi, Shuo Zhang 0003, Song Chang, Youfang Lin |
PRCV (6) | 2 |
| 2024 | Semantic-Consistency-guided Learning on Deep Features for Unsupervised Salient Object DetectionabstractUnsupervised salient object detection is an important task in many real-world scenarios where pixel-wise label information is of scarce availability. Despite its significance, this problem remains rarely explored, with a few works that consider unsupervised salient object detection methods based on the fused graph from the sum fusion of multiple deep feature similarity matrices. However, these methods ignore the interrelation of the low-level feature similarity matrices and the high-level semantic similarity matrice, which degrades the quality of the fused graph. In this article, we propose a semantic-consistency-guided multi-graph fusion learning algorithm for unsupervised saliency detection, where the consistency and inconsistency between multiple low-level feature similarity matrices and the high-level semantic similarity matrice are explored to promote the robustness and quality of the fused graph. In the first stage, a semantic-consistency-guided multi-graph fusion learning method is proposed to exploit consistency and inconsistency of multiple low-level deep features and the high-level semantic feature. The semantic-consistency-guided similarity matrices are computed for preliminary saliency ranking. In the following saliency refinement stage, the semantic-enhanced similarity matrices are built by the cross diffusion to fuse the multiple low-level deep features and the high semantic deep feature. Based on the semantic-enhanced similarity matrices, the refinement saliency maps are calculated in a semantic-enhanced cellular automata manner. Furthermore, the final ensemble stage of the large margin semi-supervised classification views the preliminary ranking results and refinement results as features, adopts the large margin graphs for saliency ensemble. Extensive evaluations over four benchmark datasets show that the proposed unsupervised method performs favorably against the state-of-the-art approaches and is competitive with some supervised deep learning-based methods. Ying Ying Zhang, Shuo Zhang 0003, Ming Hui |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2022 | Flexible Hybrid Lenses Light Field Super-Resolution using Layered RefinementabstractIn the hybrid lenses Light Field (LF) images, a high-resolution (HR) camera is in the center of the multiple low-resolution (LR) cameras, which introduces the beneficial high-frequency information for LF super-resolution. Therefore, how to effectively utilize the high-frequency information of the central view is the key issue for the hybrid lenses LF images super-resolution. In this paper, we propose a novel learning-based framework with Layered Refinement to super-resolve the hybrid lenses LF images. Specifically, we first transform the depth information of the scene into the layered position information, and refine it by complementing the high-frequency information of the HR central view to generate a high-quality representation of the depth information. Then, guided by high-quality depth representation, we propagate the information of the HR central view to the surrounding views accurately, and utilize the layered position information to maintain the occlusion relationship during the propagation. Moreover, as the generation of each layer position information is independent in our method, our trained model can flexibly adapt kinds of scenes with various disparity ranges without additional training. Experiments show that the proposed method outperforms the SOTA methods in kinds of scenes from simulated and real-world datasets with various disparity ranges. The code is available at \urlhttps://github.com/racso10/LFHSR. Song Chang, Youfang Lin, Shuo Zhang 0003 |
ACM Multimedia | 3 |
| 2021 | Attention-based Multi-Level Fusion Network for Light Field Depth EstimationabstractDepth estimation from Light Field (LF) images is a crucial basis for LF related applications. Since multiple views with abundant information are available, how to effectively fuse features of these views is a key point for accurate LF depth estimation. In this paper, we propose a novel attention-based multi-level fusion network. Combined with the four-branch structure, we design intra-branch fusion strategy and inter-branch fusion strategy to hierarchically fuse effective features from different views. By introducing the attention mechanism, features of views with less occlusions and richer textures are selected inside and between these branches to provide more effective information for depth estimation. The depth maps are finally estimated after further aggregation. Experimental results shows the proposed method achieves state-of-the-art performance in both quantitative and qualitative evaluation, which also ranks first in the commonly used HCI 4D Light Field Benchmark. Shuo Zhang 0003, Youfang Lin |
AAAI | 2 |
| 2021 | Removing Foreground Occlusions in Light Field using Micro-lens Dynamic FilterabstractForeground occlusion removal task aims to automatically detect and remove foreground occlusions and recover background objects. Since for Light Fields (LFs), background objects occluded in some views may be seen in other views, the foreground occlusion removal task for LFs is easy to achieve. In this paper, we propose a learning-based method combining ‘seeking’ and ‘generating’ to recover occluded background. Specifically, the micro-lens dynamic filters are proposed to ‘seek’ occluded background points in shifted micro-lens images and remove occlusions using angular information. The shifted images are then combined to further ‘generate’ background regions to supplement more background details using spatial information. By fully exploring the angular and spatial information in LFs, the dense and complex occlusions can be easily removed. Quantitative and qualitative experimental results show that our method outperforms other state-of-the-arts methods by a large margin. Shuo Zhang 0003, Zeqi Shen, Youfang Lin |
IJCAI | 1 |
| 2021 | Occlusion-aware Bi-directional Guided Network for Light Field Salient Object DetectionabstractExisting light field based works utilize either views or focal stacks for saliency detection. However, since depth information exists implicitly in adjacent views or different focal slices, it is difficult to exploit scene depth information from both. By comparison, Epipolar Plane Images (EPIs) provide explicit accurate scene depth and occlusion information by projected pixel lines. Due to the fact that the depth of an object is often continuous, the distribution of occlusion edges concentrates more on object boundaries compared with traditional color edges, which is more beneficial for improving accuracy and completeness of saliency detection. In this paper, we propose a learning-based network to exploit occlusion features from EPIs and integrate high-level features from the central view for accurate salient object detection. Specifically, a novel Occlusion Extraction Module is proposed to extract occlusion boundary features from horizontal and vertical EPIs. In order to naturally combine occlusion features in EPIs and high-level features in central view, we design a concise Bi-directional Guiding Flow based on cascaded decoders. The flow leverages generated salient edge predictions and salient object predictions to refine features in mutual encoding processes. Experimental results demonstrate that our approach achieves state-of-the-art performance in both segmentation accuracy and edge clarity. Dong Jing, Shuo Zhang 0003, Runmin Cong, Youfang Lin |
ACM Multimedia | 2 |
| 2021 | Enhanced Spinning Parallelogram Operator Combining Color Constraint and Histogram Integration for Robust Light Field Depth EstimationabstractAccurate depth estimation is an essential part of most light field applications. Previous approach Spinning Parallelogram Operator (SPO) achieves robust depth estimation results in noisy and occluded scenes by fitting lines in Epipolar Plane Images (EPIs). However, only using histogram distances, SPO cannot handle complex occlusions and is sensitive to different parameters. In this letter, we propose enhanced Spinning Parallelogram Operator utilizing color constraint and histogram integration (SPO-CH) to address the above problems. The color consistency between different views is first introduced to exclude occlusions when measuring the histogram distances of occluded areas. We then design a novel Gaussian integration histogram that combines information from other adjacent bins, which improves the performances of noisy and texture-less scenes. Experimental results show that the proposed method effectively improves the depth estimation results. Comparisons with other state-of-the-art approaches show that the proposed method achieves better results for both synthetic and real-world images. Weikun Wang, Youfang Lin, Shuo Zhang 0003 |
IEEE Signal Process. Lett. | 3 |
| 2021 | End-to-End Light Field Spatial Super-Resolution Network Using Multiple Epipolar GeometryabstractLight Field (LF) cameras are considered to have many potential applications since angular and spatial information is captured simultaneously. However, the limited spatial resolution has brought lots of difficulties in developing related applications and becomes the main bottleneck of LF cameras. In this paper, an end-to-end learning-based method is proposed to simultaneously reconstruct all view images in LFs with higher spatial resolution. Based on the epipolar geometry, view images in one LF are first grouped into several image stacks and fed into different network branches to learn sub-pixel details for each view image. Since LFs have dense sampling in angular domain, sub-pixel details in multiple spatial directions are learned from corresponding angular directions in multiple branches, respectively. Then, sub-pixel details from different directions are further integrated to generate global high-frequency residual details. Combined with the spatially upsampled LF, the final LF with high spatial resolution is obtained. Experimental results on synthetic and real-world datasets demonstrate that the proposed method outperforms other state-of-the-art methods in both visual and numerical evaluations. We also implement the proposed method on LFs with different angular resolution and experiments show that the proposed method achieves superior results than others, especially for LFs with small angular resolution. Furthermore, since the epipolar geometry is fully considered, the proposed network shows good performances in preserving the inherent epipolar property in LF images. Shuo Zhang 0003, Song Chang, Youfang Lin |
IEEE Trans. Image Process. | 1 |
| 2020 | Local Regression Ranking for Saliency DetectionabstractSaliency detection is an important and challenging research topic due to the variety and complex of the background and saliency regions. In this paper, we present a novel unsupervised saliency detection approach by exploiting a learning-based ranking framework. First, the local linear regression model is adopted to simulate the local manifold structure of every image element, which is approximately linear. Using the background queries from the boundary prior, we construct a unified objective function to globally minimize all the errors of the local models for the whole image element points. The Laplacian matrix is learned via optimizing the unified objective function. Low-level image features as well as high-level semantic information extracted from deep neural networks are used for the Laplacian matrix learning. Based on the learnt Laplacian matrix, the saliency of the image element is measured as the relevance ranking to the background queries. The foreground queries are obtained from the background-based saliency and the relevance ranking to the foreground queries is calculated in the same way as the background-based saliency. Second, we calculate an enhanced similarity matrix by fusing two different-level deep feature metrics through cross diffusion. A propagation algorithm uses this enhanced similarity matrix to better exploit the intrinsic relevance of similar regions and improve the saliency ranking results effectively. Results on four benchmark datasets with pixel-wise accurate labelling demonstrate that the proposed unsupervised method shows better performance compared with the newest state-of-the-art methods and is competitive with deep learning-based methods. Ying-Ying Zhang, Shuo Zhang 0003, Hai-Zhen Song, XinGang Zhang |
IEEE Trans. Image Process. | 2 |
| 2019 | Residual Networks for Light Field Image Super-ResolutionabstractLight field cameras are considered to have many potential applications since angular and spatial information is captured simultaneously. However, the limited spatial resolution has brought lots of difficulties in developing related applications and becomes the main bottleneck of light field cameras. In this paper, a learning-based method using residual convolutional networks is proposed to reconstruct light fields with higher spatial resolution. The view images in one light field are first grouped into different image stacks with consistent sub-pixel offsets and fed into different network branches to implicitly learn inherent corresponding relations. The residual information in different spatial directions is then calculated from each branch and further integrated to supplement high-frequency details for the view image. Finally, a flexible solution is proposed to super-resolve entire light field images with various angular resolutions. Experimental results on synthetic and real-world datasets demonstrate that the proposed method outperforms other state-of-the-art methods by a large margin in both visual and numerical evaluations. Furthermore, the proposed method shows good performances in preserving the inherent epipolar property in light field images. Shuo Zhang 0003, Youfang Lin, Hao Sheng 0001 |
CVPR | 1 |
| 2018 | W-Shaped Selection for Light Field Super-Resolution
Bing Su 0004, Hao Sheng 0001, Shuo Zhang 0003, Da Yang 0001, Nengcheng Chen, Wei Ke 0001 |
KSEM (1) | 3 |
| 2018 | Occlusion-aware depth estimation for light field using multi-orientation EPIs
Hao Sheng 0001, Shuo Zhang 0003, Jun Zhang 0006, Da Yang 0001 |
Pattern Recognit. | 3 |
| 2018 | Micro-Lens-Based Matching for Scene Recovery in Lenslet CamerasabstractSince a light-field camera is able to capture more information than a traditional camera, a lot of methods, such as depth estimation, image super-resolution, and view synthesis, are explored for recovering scene information. In this paper, we propose a novel framework for scene recovery based on lenslet-based light-field camera images. Instead of using traditional matching terms, we design a new micro-lens-based matching term to calculate structure information and recover several kinds of scene information simultaneously. On the one hand, inherent information in micro-lens images is selected to complement details in sub-aperture images. On the other hand, sub-aperture images are used to expand micro-lens images and synthesize new view images. A new micro-lens-based consistency metric is introduced for the matching term to handle occlusions in depth estimation and image reconstruction. The newly appeared and newly occluded areas in synthesized views are analyzed and recovered based on information from surrounding points. Experimental results show that the proposed depth estimation method outperforms state-of-the-art methods on both synthetic and lenslet-based light-field images, especially in low-texture and occlusion regions. Furthermore, the super-resolution and view synthesis methods are able to acquire view images with more details and less aliasing artifacts. Shuo Zhang 0003, Hao Sheng 0001, Da Yang 0001, Jun Zhang 0006, Zhang Xiong 0001 |
IEEE Trans. Image Process. | 1 |
| 2017 | Geometric Occlusion Analysis in Depth Estimation Using Integral Guided Filter for Light-Field ImageabstractUnlike traditional multi-view images, sampling in angular domain of light field images is distributed in different directions. Therefore, an angular sampling image (ASI), comprising of possible matching points extracted from each view, is available for each point. In this paper, we analyze the geometric relationship between ASIs and reference sub-aperture images, and then prove the occlusion boundary similarity. Based on the geometric relationship in extreme cases, we show that some points in ASI have higher reliability than other points for depth calculation. An integral guided filter is then built based on the sub-aperture image to predict occlusion probabilities in ASIs. The filter is independent of ASIs and has no requirement for high angular resolution so that it is easy to apply to the cost volume calculation. We integrate the filter into our depth estimation framework and other state-of-the-art depth estimation frameworks. Experimental results demonstrate that the proposed filter is more effective to occluded point detection in ASIs than other methods. Results from different data sets show that our method outperforms the existing state-of-the-art depth estimation methods, especially along occlusion boundaries. Hao Sheng 0001, Shuo Zhang 0003, Xiaochun Cao, Yajun Fang, Zhang Xiong 0001 |
IEEE Trans. Image Process. | 2 |
| 2016 | Saliency analysis based on depth contrast increasedabstractHumans can understand their surroundings by an additional depth cue that provides by stereopsis, which plays an important role in the human visual system. Recently, depth saliency has been attracted much attention. But depth image differs a lot from color image. Feature extraction in depth image is an important problem in depth saliency analysis. Previous studies always extract features from depth map directly. This paper proposes a method which can make the saliency analysis easier and more accurate by increasing the depth contrast between the salient object and distractors. Then, we extended a recent saliency analysis approach to evaluate the saliency of the difference map. Finally, after the optimization by depth information and color information the final saliency map can be obtained. Our experiments on public dataset show that our method significantly outperforms state-of-the-art. Hao Sheng 0001, Shuo Zhang 0003 |
ICASSP | 3 |
| 2016 | Relative location for light field saliency detectionabstractLight field images, which capture multiple images from different angles of a scene, have been proved that can detect salient regions more effectively. Instead of estimating depth labels from light field images, we proposed to extract relative locations, which can distinguish whether the object is located before the focus plane of the main lens or not, for saliency detection. The relative locations are calculated by comparing raw light field images captured by plenoptic cameras and central views of scenes. The relative locations are then integrated to a modified saliency detection framework to obtain the salient regions. Experimental results demonstrate that the proposed relative locations can help to improve the accuracy of results, which is also efficient. Moreover, the modified framework outperforms the state-of-the-art methods for light field images saliency detection. Hao Sheng 0001, Shuo Zhang 0003, Zhang Xiong 0001 |
ICASSP | 2 |
| 2016 | Segmentation of light field image with the structure tensorabstractWe propose a segmentation model for light field images based on superpixels segmentation and graph-cuts algorithm. Unlike traditional images, which do not offer information for different directions, a light field image encodes space data which can be computed on its epipolar plane images (EPI) with some effective methods. In our work, we analyze the structure of EPI and research the computational process of disparity using EPI. On this basis, we present a new method for computing disparity using the modified structure tensor on EPIs. We further apply the computed disparity labels by fusing RGB images and disparity labels to obtain more detailed over-segmentation. Meanwhile, the modified structure tensor algorithm is used to get more accurate image boundaries, which plays a role in computing disparity features. All these processes are applied in an interactive segmentation model. Our experiments on public data sets demonstrate that the proposed light field image segmentation achieves a higher performance compared with state-of-the-art methods. Hao Sheng 0001, Senyou Deng, Shuo Zhang 0003, Chao Li 0001, Zhang Xiong 0001 |
ICIP | 3 |
| 2016 | Cellular Automata Based on Occlusion Relationship for Saliency Detection
Hao Sheng 0001, Weichao Feng, Shuo Zhang 0003 |
KSEM | 3 |
| 2016 | Robust depth estimation for light field via spinning parallelogram operator
Shuo Zhang 0003, Hao Sheng 0001, Chao Li 0001, Jun Zhang 0006, Zhang Xiong 0001 |
Comput. Vis. Image Underst. | 1 |
| 2015 | Guided integral filter for light field stereo matchingabstractDifferent from the traditional multi-view images, the sampling in angular resolution of light field images is continuous in each direction. Therefore, an angular sampling image, comprising of matching points extracted from each view, can be constructed at different possible depth for each point. In this paper, we prove that the angular sampling image of an occluded point at the correct depth is similar to the scene around the point. On this basis, a guided integral filter acquired from the reference view is proposed to weight the matching points for consistency measure, which predicts the probabilities of occlusions. The discrete labeling problem is then solved by a filter-based algorithm to approximate the optimal solution. Experimental results demonstrate that the proposed method outperforms the state-of-the-art methods for light field depth estimation both in occluded and ambiguous regions, and has no requirement for angular resolution. Hao Sheng 0001, Shuo Zhang 0003, Gengliang Zhu, Zhang Xiong 0001 |
ICIP | 2 |