Ruixuan Cong

dblp:332/7157 · DBLP profile ↗
← Back
33ranked-venue papers
7as first author
33since 2021 · last 2026
0000-0001-6410-5248ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 16 · 5 first-author · 16 since 2021Artificial intelligence and machine learning · 14 · 3 first-author · 14 since 2021Systems, architecture and hardware · 3 · 3 since 2021Security and privacy · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Difference-guided full-view volume for light field depth estimation
Tun Wang, Hao Sheng 0001, Ruixuan Cong, Da Yang 0001, Zhenglong Cui, Guanqun Su
Expert Syst. Appl.3
2026 Learning Three-Domain Implicit Image Function for Arbitrary-Scale Light Field Super-Resolution
abstract
Various deep learning-based light field image super-resolution methods have attained notable success in recent years. However, most of them focus on encoder design while neglecting the critical role of upsampling process in decoder part. Motivated by the recent progress in single image domain with implicit neural representation, we elaborately propose a spatial-angular-epipolar implicit image function (SAEIIF) in this paper, which can redefine the upsampling process to significantly improve performance and enable arbitrary-scale light field super-resolution. Specifically, it contains two complementary upsampling branches. One branch incorporates spatial implicit image function (SIIF) and angular implicit image function (AIIF) to mine intra-view information in sub-aperture images and inter-view information in macro pixels. The other branch involves epipolar implicit image function (EIIF) to leverage spatial-angular correlation in epipolar plane images. By decomposing SIIF, AIIF and EIIF into horizontal and vertical two-step upsampling to form a perfect match of upsampling scale, SAEIIF introduces a multi-stage feature interaction architecture across two branches to fully merge spatial, angular and epipolar domain information. Furthermore, we optimize feature sampling strategy based on characteristics of sub-aperture images, macro pixels, and epipolar plane images, introducing horizontal-vertical separable local sampling for SIIF and AIIF, as well as dual-source oriented line sampling used for EIIF. The extensive experimental results demonstrate that our SAEIIF can be effectively integrated with most encoders and achieve outstanding performance on both fixed-scale and arbitrary-scale light field spatial super-resolution, angular super-resolution, spatial-angular joint super-resolution.
Ruixuan Cong, Hao Sheng 0001, Da Yang 0001, Zhenglong Cui, Weifeng Lyv, Wei Ke 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
2026 Learn depth space from light field via a distance-constraint query mechanism
abstract
The Light Field (LF) captures both spatial and angular information of scenes, enabling precise depth estimation. Recent advancements in deep learning have led to significant success in this field; however, existing methods primarily focus on modeling surface characteristics (e.g., depth maps) while overlooking the depth space, which contains additional valuable information. The depth space consists of numerous space points and provides substantially more geometric data than a single depth map. In this paper, we conceptualize depth prediction as a spatial modeling problem, aiming to learn the entire depth space rather than merely a single depth map. Specifically, we define space points as signed distances relative to the scene surface and propose a novel distance-constraint query mechanism for LF depth estimation. To model the depth space effectively, we first develop a mixed sampling strategy to approximate its data representation. Subsequently, we introduce an encoder-decoder network architecture to query the distances of each point, thereby implicitly embedding the depth space. Finally, to extract the target depth map from this space, we present a generation algorithm that iteratively invokes the decoder network. Through extensive experiments, our approach achieves the highest performance on LF depth estimation benchmarks, and also demonstrates superior performance on various synthetic and real-world scenes.
Hao Sheng 0001, Rongshan Chen, Ruixuan Cong, Da Yang 0001, Zhenglong Cui
Pattern Recognit.3
2026 A geometry-aware implicit-explicit framework for light field angular super-resolution
Mingyuan Zhao 0001, Da Yang 0001, Rongshan Chen, Zhenglong Cui, Ruixuan Cong, Hao Sheng 0001
Pattern Recognit.5
2026 Gradient-Guided Density Redistribution Network for Light Field Full-View Depth Estimation
abstract
Light field (LF) full-view depth estimation aims to recover dense and coherent depth maps for all sub-aperture views, which is crucial for applications such as 3D reconstruction, LF editing and virtual reality. However, directly extending center-view volume-based methods to the full-view is computationally infeasible, as it requires constructing a separate cost volume for each view. Besides, existing full-view propagation-based approaches, while more efficient, frequently suffer from edge fattening and cross-view inconsistencies in the presence of occlusions. In this paper, we propose a gradient-guided density redistribution network (GDRNet), a novel end-to-end framework that efficiently generates full-view depth maps by constructing a single plane-density volume and a multi-plane depth image, which are then propagated to all angular views. To resolve ambiguous estimates at occlusion edges, we perform a direction-aware gradient-guided density redistribution only inside a dilated edge narrow band. For each center pixel in edge regions, a guidance gradient is derived from the initial depth map to determine the normal and tangent directions. Then, density in edge fattening regions can be redistributed via sampling along the normal direction, while similarity along the tangent direction can fill bad pixels with inconsistencies. Furthermore, an adaptive edge extraction module with four directional learnable Sobel kernels is designed to jointly exploit spatial and angular gradients, enabling robust detection and localizing the refinement band. Extensive experiments on synthetic and real-world LF datasets demonstrate that GDRNet achieves state-of-the-art accuracy and edge sharpness in both quantitative and qualitative evaluations, while maintaining computational efficiency compared to full-view methods.
Tun Wang, Zhenglong Cui, Ruixuan Cong, Da Yang 0001, Mingyuan Zhao 0001, Hao Sheng 0001
IEEE Trans. Circuits Syst. Video Technol.3
2026 Combining Gaussian and Pixel Representation for Light Field View Reconstruction
abstract
Light field (LF) benefits various applications due to its rich spatial and angular information. To address the technical limitation in terms of imaging resolution, LF view reconstruction becomes a research hotspot. However, relevant methods mainly focus on pixel representation modeling on image plane but ignore the importance of scene geometry modeling. Inspired by powerful geometry description ability embedded in 3D Gaussian Splatting, we construct a network called LFGaussian to perform generalizable LF view reconstruction in this paper. Specifically, owing to the unique composition of cross-view Gaussian attribute deviation under 4D LF imaging setting, we propose disparity-guided feed-forward 2D Gaussian propagation with novel Gaussian primitive definition, subtly implementing Gaussian unprojection-projection operation in camera parameter-free case. On this basis, we introduce a dual-branch workflow including Gaussian representation rendering and pixel representation upsampling to create features of target views from two different levels, which complement each other to jointly realize geometric structure consistency as well as texture detail consistency across all target views. Besides, for the pursuit of high-efficient and high-quality Gaussian representation rendering, we design sub-sampling Gaussian decoding to alleviate Gaussian redundancy and leverage Gaussian splitting to allocate additional Gaussians for complex geometry regions identified by disparity gradient. Experimental results show that the proposed LFGaussian achieves superior performance compared with state-of-the-art methods on both real-world and synthetic LF datasets, proving the effectiveness of introducing Gaussian representation for LF view reconstruction. Furthermore, our LFGaussian supports arbitrary-scale reconstruction, showing high flexibility for the upsampling scale factor.
Ruixuan Cong, Zhenglong Cui, Wei Ke 0001, Weifeng Lyv, Hao Sheng 0001
IEEE Trans. Image Process.1
2026 UCGR: Closing the Discretization Gap in Light Field Depth Estimation via Unified Continuous Geometry Representation
abstract
Light field (LF) cameras encode dense spatial-angular information for depth estimation, critical for applications such as 3D reconstruction, refocusing, and virtual reality. However, current deep learning methods for LF depth estimation still face significant challenges due to the discretization gap between the continuous geometry of real scenes and the discrete sampling of digital images, limiting their effectiveness in high-precision application. This gap manifests in two complementary forms: spatial discretization leads to structural ambiguities and information loss, while depth discretization introduces inaccuracies due to fixed, discrete depth sampling. To address these challenges, we propose Unified Continuous Geometry Representation (UCGR), a unified representation that models scene geometry as a continuous field over image coordinates and depth. UCGR treats spatial and depth discretization as two facets of the same problem and realized by two complementary operators: (1) Adaptive Plane Sampling Operator, which learns edge-aware planar priors to preserve geometric details and mitigate spatial discretization. (2) Contextual Depth Correction Operator, which utilizes contextual information for depth correction, ensuring continuous depth estimation and suppressing artifacts. Building on UCGR, we propose a Continuous Geometry Network (CGNet) that collaboratively optimizes both spatial and depth discretization for accurate and consistent LF depth estimation. Extensive experiments on synthetic and real-world LF datasets demonstrate that CGNet achieves state-of-the-art performance, significantly outperforming existing LF depth estimation methods in terms of both accuracy and robustness.
Zexin Sun, Tun Wang, Rongshan Chen, Ruixuan Cong, Wei Ke 0001, Hao Sheng 0001
IEEE Trans. Vis. Comput. Graph.4
2025 Rethinking the Upsampling Process in Light Field Super-Resolution with Spatial-Epipolar Implicit Image Function
Ruixuan Cong, Mingyuan Zhao 0001, Da Yang 0001, Rongshan Chen, Hao Sheng 0001
ICCV1
2025 Stereo matching on epipolar plane image for light field depth estimation via oriented structure
Rongshan Chen, Hao Sheng 0001, Ruixuan Cong, Da Yang 0001, Zhenglong Cui, Wei Ke 0001
Eng. Appl. Artif. Intell.3
2025 Four-dimension efficient pixel-frequency transformer for light field spatial and angular super-resolution
Hao Sheng 0001, Ruixuan Cong, Da Yang 0001, Rongshan Chen, Zhenglong Cui
Eng. Appl. Artif. Intell.3
2025 Progressive epipolar geometry for robust light field super-resolution
Hao Zhang 0146, Hao Sheng 0001, Rongshan Chen, Da Yang 0001, Ruixuan Cong, Zhenglong Cui, Xuefei Huang, Guanqun Su
Eng. Appl. Artif. Intell.5
2025 Pixel-wise matching cost function for robust light field depth estimation
Rongshan Chen, Hao Sheng 0001, Da Yang 0001, Zhenglong Cui, Ruixuan Cong
Expert Syst. Appl.6
2025 Multiplane depth image for view-consistent light field depth estimation
Tun Wang, Hao Sheng 0001, Rongshan Chen, Ruixuan Cong, Mingyuan Zhao 0001, Da Yang 0001
Knowl. Based Syst.4
2025 A GPU-Enabled Framework for Light Field Efficient Compression and Real-Time Rendering
abstract
Real-time rendering offers instantaneous visual feedback, making it crucial for mixed-reality applications. The light field captures both light intensity and direction in a 3D environment, serving as a data-rich medium to enhance mixed-reality experiences. However, two major challenges remain: 1) current light field rendering techniques are unsuitable for real-time computation, and 2) existing real-time methods cannot efficiently process high-dimensional light field data on GPU platforms. To overcome these challenges, we propose an framework utilizing a compact neural representation of light field data, implemented on a GPU platform for real-time rendering. This framework provides both compact storage and high-fidelity real-time computation. Specifically, we introduce a ray global alignment strategy to simplify the framework and improve practicality. This strategy enables the learning of an optimal embedding for all local rays in a globally consistent way, removing the need for camera pose calculations. To achieve effective compression, the neural light field is employed to map each embedded ray to its corresponding color. To enable real-time rendering, we design a novel super-resolution network to enhance rendering speed. Extensive experiments demonstrate that our framework significantly enhances compression efficiency and real-time rendering performance, achieving nearly 50$\mathbf{\times}$compression ratio and 100 FPS rendering.
Mingyuan Zhao 0001, Hao Sheng 0001, Rongshan Chen, Ruixuan Cong, Tun Wang, Zhenglong Cui, Da Yang 0001, Shuai Wang 0027, Wei Ke 0001
IEEE Trans. Computers4
2025 Surface-Continuous Scene Representation for Light Field Depth Estimation via Planarity Prior
abstract
Light field (LF) imaging captures both spatial and angular information of the real world, enabling precise depth estimation. However, images are merely discrete expressions of scenes. Limited by imaging technology, LF camera cannot capture the infinite rays emitted by scenes, leading to the discrete information storage (e.g. pixel). Consequently, previous deep learning methods have encountered challenges in accurately extracting depth information from LF images. In this paper, we investigate a surface-continuous scene representation using planarity prior and design PlaneNet, a Plane-based Network that successfully generates highly detailed depth maps for real scenes. Specifically, inspired by the plane assumption that real-world scenes generally yield piecewise smooth surfaces, we refine it to the pixel level for continuous surface approximation, which can overcome the limitations of discrete representation. Rather than explicitly parameterizing planes as multiple coefficients, we propose a novel plane regular sampling operator (PRSO), enabling the network to fit smooth depth surfaces easily. To explore the role of our theory at the feature level, we also introduce PRSO into the intermediate layers of PlaneNet. Experiments show that our method achieves state-of-the-art performance on both synthetic and real-world LF scenes, ranking 1st (MSE) on the HCI 4D Light Field benchmark. Furthermore, we explore the utilization of our representation in multiple LF depth estimation networks, and experiments demonstrate improved performance when surface-continuous representation is applied. Code is available athttps://github.com/crs904620522/PlaneNet.
Rongshan Chen, Hao Sheng 0001, Da Yang 0001, Zhenglong Cui, Ruixuan Cong
IEEE Trans. Circuits Syst. Video Technol.5
2025 UNeLF: Unconstrained Neural Light Field for Self-Supervised Angular Super-Resolution
abstract
Compared to supervised learning methods, self-supervised learning methods address the domain gap problem between light field (LF) datasets collected under varying acquisition conditions, which typically leads to decreased performance when differences exist in the distribution between the training and test sets. However, current self-supervised light field angular super-resolution (LFASR) techniques primarily focus on exploiting discrete spatial-angular features while neglecting continuous LF information. In contrast to previous work, we propose a self-supervised unconstrained neural light field (UNeLF) to continuously represent LF for LFASR. Specifically, any LF can be described as the camera pose for each sub-aperture image (SAI) and the two-plane that captures these SAIs. To describe the former, we introduce a SAIs-dependent pose optimization method to solve the issue that arises from the narrow baseline of most LF data, which hinders robust camera pose estimation. This mechanism reduces the number of trainable camera parameters from a quadratic to a constant scale, thereby alleviating the complexity of joint optimization. For the latter, we propose a novel adaptive two-plane parameterization strategy to determine the two-plane that captures these SAIs, facilitating refocusing. Finally, we jointly optimize the camera parameters, near-far planes and neural light field, efficiently mapping each adaptive two-plane parameterized ray to its correspondence color in a continuous manner. Comprehensive experiments demonstrate that UNeLF achieves faster training and inference with fewer computational resources while exhibiting superior performance on both synthetic and real-world datasets.
Mingyuan Zhao 0001, Hao Sheng 0001, Rongshan Chen, Ruixuan Cong, Zhenglong Cui, Da Yang 0001
IEEE Trans. Circuits Syst. Video Technol.4
2025 Towards Depth-Continuous Scene Representation With a Displacement Field for Robust Light Field Depth Estimation
abstract
Light field (LF) captures both spatial and angular information of scenes, enabling accurate depth estimation. However, previous deep learning methods have typically model surface depth only, while ignoring the continuous nature of depth in 3D scenes. In this paper, we use displacement field (DF) to describe this continuous property, and propose a novel depth-continuous scene representation for robust LF depth estimation. Experiments demonstrate that our representation enables the network to generate highly detailed depth maps with fewer parameters and faster speed. Specifically, inspired by signed distance field in 3D object description, we aim to exploit the intrinsic depth-continuous property of 3D scenes using DF, and define a novel depth-continuous scene representation. Then, we introduce a simple yet general learning framework for depth-continuous scene embedding, and the proposed network, DepthDF, achieves state-of-the-art performance on both synthetic and real-world LF datasets, ranking 1st on the HCI 4D Light Field benchmark. Furthermore, previous LF depth estimation methods can also be seamlessly integrated into this framework. Finally, we extend this framework beyond LF depth estimation to various tasks, including multi-view stereo depth inference, LF super-resolution, and LF salient object detection. Experiments demonstrate improved performance when the continuous scene representation is applied, suggesting that our framework can potentially bring insights to more fields.
Rongshan Chen, Hao Sheng 0001, Da Yang 0001, Ruixuan Cong, Zhenglong Cui, Tun Wang, Mingyuan Zhao 0001
IEEE Trans. Multim.4
2025 Multiplane-Based Cross-View Interaction Mechanism for Robust Light Field Angular Super-Resolution
abstract
Dense sampling of the light field (LF) is essential for various applications, such as virtual reality. However, the collection process is prohibitively expensive due to technological limitations in imaging. Synthesizing novel views from sparse LF data, known as LF Angular Super-Resolution (LFASR), offers an effective solution to this problem. Accurate cross-view interaction is crucial for this task, given the complementary information between LF views. Previous methods, however, suffer from limited reconstruction quality due to inefficient view interaction. To address this, we propose a Multiplane-based Cross-view Interaction Mechanism (MCIM) for robust LFASR. Extensive comparisons with state-of-the-art methods demonstrate that our method achieves superior performance, both visually and quantitatively. Specifically, Drawing inspiration from MultiPlane Images (MPI) in scene modeling, our mechanism incorporates a novel Multiplane Feature Fusion (MPFF) strategy. This strategy facilitates fast and accurate cross-view interaction, enhancing the network's robustness to scene geometry and suitability for different-baseline LF scenes. Furthermore, to address information redundancy in multiplanes, we leverage the transparency property of MPI and devise a plane selection strategy. Finally, we propose CSTNet, a Cross-Shaped Transformer-based network for LFASR, which employs a cross-shaped self-attention mechanism to enable low-cost training and inference. Experimental results on various angular super-resolution tasks validate that our network achieves state-of-the-art performance on both synthetic and real-world LF scenes.
Rongshan Chen, Hao Sheng 0001, Da Yang 0001, Ruixuan Cong, Zhenglong Cui
IEEE Trans. Vis. Comput. Graph.4
2025 View-Guided Cost Volume for Light Field Arbitrary-View Disparity Estimation
abstract
Per-view disparity estimation for light field (LF) is critical for various applications such as light field editing, but previous work mostly focuses on estimating disparity for the center view. In this paper, we propose a view-guided cost volume (VGCV), which successfully generates high-quality disparity maps for LF arbitrary view. Unlike previous methods that construct a static cost for center view only, VGCV is designed with view information and can be applicable to arbitrary-view estimation. In particular, since the key to achieving it is to condition cost on view, we extend previous static cost to a conditional one by introducing the spatial and angular information of target view into cost construction and aggregation, experiments show that this way can effectively adapt VGCV to arbitrary-view task. For construction, previous stereo-matching methods usually adopt correlation (e.g., variance) for dynamic estimation, but just using correlation can lose image structure information, which is essential for scene detail recovery, therefore we design an image-guided construction module and use cross-view attention to adapt cost for conditional construction while keeping its spatial information. Then for aggregation, we present a coordinate-guided aggregation module for VGCV regularization, which is specially designed to solve the problem of LF view deviation. Finally, we implement a Light Field Arbitrary-View Disparity Estimation Network (LFAVNet), then perform it on both synthetic and real LFs. Experiments demonstrate that LFAVNet can generate a higher-quality disparity map for arbitrary view in LF. We also extend our method to center-view estimation and light field editing tasks, which all achieve advanced performance.
Rongshan Chen, Hao Sheng 0001, Da Yang 0001, Zhenglong Cui, Ruixuan Cong, Shuai Wang 0027
IEEE Trans. Vis. Comput. Graph.6
2024 Multi-scale Spatial-Angular Information Aggregation Network for Image Semantic Segmentation
Ruixuan Cong, Hao Sheng 0001
ICONIP (8)2
2024 Light field depth estimation: A comprehensive survey from principles to future
abstract
Light field (LF) depth estimation is an important research direction in the area of computer vision and computational photography, which aims to infer the depth information of different objects in three-dimensional scenes by capturing LF data. Given this new era of significance, this article introduces a survey of the key concepts, methods, novel applications, and future trends in this area. We summarize the LF depth estimation methods, which are usually based on the interaction of radiance from rays in all directions of the LF data, such as epipolar-plane, multi-view geometry, focal stack, and deep learning. We analyze the many challenges facing each of these approaches, including complex algorithms, large amounts of computation, and speed requirements. In addition, this survey summarizes most of the currently available methods, conducts some comparative experiments, discusses the results, and investigates the novel directions in LF depth estimation.
Tun Wang, Hao Sheng 0001, Rongshan Chen, Da Yang 0001, Zhenglong Cui, Ruixuan Cong, Mingyuan Zhao 0001
High Confid. Comput.7
2024 A survey for light field super-resolution
abstract
Compared to 2D imaging data, the 4D light field (LF) data retains richer scene’s structure information, which can significantly improve the computer’s perception capability, including depth estimation, semantic segmentation, and LF rendering. However, there is a contradiction between spatial and angular resolution during the LF image acquisition period. To overcome the above problem, researchers have gradually focused on the light field super-resolution (LFSR). In the traditional solutions, researchers achieved the LFSR based on various optimization frameworks, such as Bayesian and Gaussian models. Deep learning-based methods are more popular than conventional methods because they have better performance and more robust generalization capabilities. In this paper, the present approach can mainly divided into conventional methods and deep learning-based methods. We discuss these two branches in light field spatial super-resolution (LFSSR), light field angular super-resolution (LFASR), and light field spatial and angular super-resolution (LFSASR) , respectively. Subsequently, this paper also introduces the primary public datasets and analyzes the performance of the prevalent approaches on these datasets. Finally, we discuss the potential innovations of the LFSR to propose the progress of our research field.
Mingyuan Zhao 0001, Hao Sheng 0001, Da Yang 0001, Ruixuan Cong, Zhenglong Cui, Rongshan Chen, Tun Wang, Shuai Wang 0027
High Confid. Comput.5
2024 Blockchain-Based Distributed Multiagent Reinforcement Learning for Collaborative Multiobject Tracking Framework
abstract
With the development of smart cities, video surveillance has become more prevalent in urban areas. The rapid growth of data brings challenges to video processing and analysis. Multi-object tracking (MOT), one of the most fundamental tasks in computer vision, has a wide range of applications and development prospects. MOT aims to locate multiple objects and maintain their unique identities by analyzing the video frame by frame. Most existing MOT frameworks are deployed in centralized systems, which are convenient for management but have problems such as weak algorithm adaptability, limited system scalability, and poor data security. In this paper, we propose a distributed MOT algorithm based on multi-agent reinforcement learning (DMARL-Tracker), which formulates MOT as a Markov decision process (MDP). Each object adjusts its tracking strategy during interactions with the environment. The benchmark results on MOT17 and MOT20 prove that our proposed algorithm achieves state-of-the-art (SOTA) performance. Based on this, we further integrate DMARL-Tracker into the blockchain and propose a blockchain-based collaborative MOT framework. All nodes collaborate and share information through the blockchain, achieving adaptation in different complex scenarios while ensuring data security. The simulation results show that our framework achieves good performance in terms of tracking and resource consumption.
Hao Sheng 0001, Shuai Wang 0027, Ruixuan Cong, Da Yang 0001, Yang Zhang 0032
IEEE Trans. Computers4
2024 An Occlusion and Noise-Aware Stereo Framework Based on Light Field Imaging for Robust Disparity Estimation
abstract
Stereo vision is widely studied for depth information extraction. However, occlusion and noise pose significant challenges to traditional methods due to failure in photo consistency. In this paper, an occlusion and noise-aware stereo framework named ONAF is proposed to get a robust depth estimation by integrating the advantages of correspondence cues and refocusing cues from light field(LF). ONAF consists of two special depth cue extractors: correspondence depth cue extractor (CCE) and refocusing depth cue extractor (RCE). CCE extracts accurate correspondence depth cues in occlusion areas based on multi-direction Ray-Epipolar Plane Images(Ray-EPIs) from LF, which are more robust than traditional multi-direction EPIs. RCE generates accurate refocusing depth cues in noise areas, benefitting from the many-to-one integration strategy and the directional perception of texture and occlusion based on multi-direction focal stacks from LF. Attention mechanism is introduced to complementarily fuse CCE and RCE to generate optimum depth maps. The experimental results prove the effectiveness of ONAF, which outperforms state-of-the-art disparity estimation methods, especially in occlusion and noise areas.
Da Yang 0001, Zhenglong Cui, Hao Sheng 0001, Rongshan Chen, Ruixuan Cong, Shuai Wang 0027, Zhang Xiong 0001
IEEE Trans. Computers5
2024 End-to-End Semantic Segmentation Utilizing Multi-Scale Baseline Light Field
abstract
Semantic segmentation based on 4D light field (LF) images exhibits superior performance by exploiting rich spatial and angular information. However, current methods only focus on narrow-baseline cases, ignoring the feasibility and capability of large disparity scene for segmentation. Motivated by this, we propose a novel network called LF-IENet++ suitable for both narrow-baseline LF and wide-baseline LF in this paper, which fully mines complementary information across views via implicit feature integration and explicit feature propagation. In order to concentrate on inconsistent context between view images during feature integration, we shield small disparity regions manifested as repeat content to avoid redundant attention. Besides, a two-stage operation consisting of the image-level warping and feature-level warping is introduced to mitigate the propagation distortion. Since both feature integration and feature propagation require exact guidance from prior disparity, we design a semantic-aware disparity estimator that leverages semantic cues to optimize disparity generation while ensuring that our network can perform semantic segmentation in an end-to-end solution. To validate the effectiveness of the proposed method, we present the first multi-scale baseline dataset for LF semantic segmentation. Compared to state-of-the-art methods, our LF-IENet++ achieves outstanding performance and shows high robustness under different disparity situations. Besides, our method obtains higher accuracy on wide-baseline cases, demonstrating the significance of introducing large disparity LF for semantic segmentation.
Ruixuan Cong, Hao Sheng 0001, Dazhi Yang 0003, Da Yang 0001, Rongshan Chen, Zhenglong Cui
IEEE Trans. Circuits Syst. Video Technol.1
2024 Multimodal Perception Integrating Point Cloud and Light Field for Ship Autonomous Driving
abstract
Robust scene perception is an essential prerequisite to ensure the reliability in ship autonomous driving. However, it is a challenging task in inland river because of the complicated and changeable environment as well as high-density ships in narrow waterway. As one of the primary technologies, obstacle trajectory locating and tracking has been widely explored in recent years. Current approaches strictly rely on lidar as only depth awareness sensor and the limited measurement range severely restricts them for distant object identification. On this account, we creatively propose a point cloud-light field fusion perception framework in this paper for the first time. Specifically, in detection stage, the former undertakes precise close object perception and the latter completes distant object locating through light field stereo matching. In tracking stage, a novel four-phase data association that combines multiple attributes from position, point cloud and image domains is utilized for accurate object matching across frames. To validate the effectiveness of our multimodal perception strategy, we implement an acquisition system consisting of two lidars and four sets of simplified light field cameras on a ship to conduct actual testing. Extensive experimental results show that the proposed framework achieves superior 3D object locating and tracking performance, far surpassing the state-of-the-art methods in terms of accuracy and real-time.
Ruixuan Cong, Hao Sheng 0001, Mingyuan Zhao 0001, Dazhi Yang 0003, Tun Wang, Rongshan Chen
IEEE Trans. Intell. Transp. Syst.1
2024 Exploiting Spatial and Angular Correlations With Deep Efficient Transformers for Light Field Image Super-Resolution
abstract
Global context information is particularly important for comprehensive scene understanding. It helps clarify local confusions and smooth predictions to achieve fine-grained and coherent results. However, most existing light field processing methods leverage convolution layers to model spatial and angular information. The limited receptive field restricts them to learn long-range dependency in LF structure. In this article, we propose a novel network based on deep efficient transformers (i.e.,LF-DET) for LF spatial super-resolution. It develops a spatial-angular separable transformer encoder with two modeling strategies termed as sub-sampling spatial modeling and multi-scale angular modeling for global context interaction. Specifically, the former utilizes a sub-sampling convolution layer to alleviate the problem of huge computational cost when capturing spatial information within each sub-aperture image. In this way, our model can cascade more transformers to continuously enhance feature representation with limited resources. The latter processes multi-scale macro-pixel regions to extract and aggregate angular features focusing on different disparity ranges to well adapt to disparity variations. Besides, we capture strong similarities among surrounding pixels by dynamic positional encodings to fill the gap of transformers that lack of local information interaction. The experimental results on both real-world and synthetic LF datasets confirm our LF-DET achieves a significant performance improvement compared with state-of-the-art methods. Furthermore, our LF-DET shows high robustness to disparity variations through the proposed multi-scale angular modeling.
Ruixuan Cong, Hao Sheng 0001, Da Yang 0001, Zhenglong Cui, Rongshan Chen
IEEE Trans. Multim.1
2024 Triple Consistency for Transparent Cheating Problem in Light Field Depth Estimation
abstract
Depth estimation extracting scenes' structural information is a key step in various light field(LF) applications. However, most existing depth estimation methods are based on the Lambertian assumption, which limits the application in non-Lambertian scenes. In this paper, we discover a unique transparent cheating problem for non-Lambertian scenes which can effectively spoof depth estimation algorithms based on photo consistency. It arises because the spatial consistency and the linear structure superimposed on the epipolar plane image form new spurious lines. Therefore, we propose centrifugal consistency and centripetal consistency for separating the depth information of multi-layer scenes and correcting the error due to the transparent cheating problem, respectively. By comparing the distributional characteristics and the number of minimal values of photo consistency and centrifugal consistency, non-Lambertian regions can be efficiently identified and initial depth estimates obtained. Then centripetal consistency is exploited to reject the projection from different layers and to address transparent cheating. By assigning decreasing weights radiating outward from the central view, pixels with a concentration of colors close to the central viewpoint are considered more significant. The problem of underestimating the depth of background caused by transparent cheating is effectively solved and corrected. Experiments on synthetic and real-world data show that our method can produce high-quality depth estimation under the transparency and the reflectivity of 90% to 20%. The proposed triple-consistency-based algorithm outperforms state-of-the-art LF depth estimation methods in terms of accuracy and robustness.
Zhenglong Cui, Da Yang 0001, Hao Sheng 0001, Rongshan Chen, Ruixuan Cong, Wei Ke 0001
IEEE Trans. Multim.6
2023 Take Your Model Further: A General Post-refinement Network for Light Field Disparity Estimation via BadPix Correction
abstract
Most existing light field (LF) disparity estimation algorithms focus on handling occlusion, texture-less or other areas that harm LF structure to improve accuracy, while ignoring other potential modeling ideas. In this paper, we propose a novel idea called Bad Pixel (BadPix) correction for method modeling, then implement a general post-refinement network for LF disparity estimation: Bad-pixel Correction Network (BpCNet). Given an initial disparity map generated by a specific algorithm, we assume that all BadPixs on it are in a small range. Then BpCNet is modeled as a fine-grained search strategy, and a more accurate result can be obtained by evaluating the consistency of LF images in this limited range. Due to the assumption and the consistency between input and output, BpCNet can perform as a general post-refinement network, and can work on almost all existing algorithms iteratively. We demonstrate the feasibility of our theory through extensive experiments, and achieve remarkable performance on the HCI 4D Light Field Benchmark.
Rongshan Chen, Hao Sheng 0001, Da Yang 0001, Zhenglong Cui, Ruixuan Cong
AAAI6
2023 Combining Implicit-Explicit View Correlation for Light Field Semantic Segmentation
abstract
Since light field simultaneously records spatial information and angular information of light rays, it is considered to be beneficial for many potential applications, and semantic segmentation is one of them. The regular variation of image information across views facilitates a comprehensive scene understanding. However, in the case of limited memory, the high-dimensional property of light field makes the problem more intractable than generic semantic segmentation, manifested in the difficulty of fully exploiting the relationships among views while maintaining contextual information in single view. In this paper, we propose a novel network called LF-IENet for light field semantic segmentation. It contains two different manners to mine complementary information from surrounding views to segment central view. One is implicit feature integration that leverages attention mechanism to compute inter-view and intra-view similarity to modulate features of central view. The other is explicit feature propagation that directly warps features of other views to central view under the guidance of disparity. They complement each other and jointly realize complementary information fusion across views in light field. The proposed method achieves outperforming performance on both real-world and synthetic light field datasets, demonstrating the effectiveness of this new architecture.
Ruixuan Cong, Da Yang 0001, Rongshan Chen, Zhenglong Cui, Hao Sheng 0001
CVPR1
2023 MFSRNet: spatial-angular correlation retaining for light field super-resolution
Hao Sheng 0001, Da Yang 0001, Zhenglong Cui, Ruixuan Cong, Wei Ke 0001
Appl. Intell.5
2023 Cross-View Recurrence-Based Self-Supervised Super-Resolution of Light Field
abstract
Compared with external-supervised learning-based (ESLB) methods, self-supervised learning-based (SSLB) methods can overcome the domain gap problem caused by different light field (LF) acquisition conditions, which results in the performance degradation of light field super-resolution on unseen test datasets. Current SSLB methods exploit the cross-scale recurrence feature in the single view image for super-resolution, ignoring the correlation information among views. Different from previous works, we propose a cross-view recurrence-based self-supervised mapping framework to correlate complementary information among views in the down-scaled input LF. Specifically, the cross-view recurrence information consists of geometry structure features and similar structure features. The former is to provide sub-pixel information according to disparity correlations among adjacent views, and the latter is to acquire similar color and contour information among arbitrary views, which can compensate for error disparity guidance of geometry structure features in sharp variance areas. Moreover, instead of the widely used “All-to-All” strategy, we propose a “Part-to-Part” mapping strategy, which is better competent for SSLB approaches with limited training examples solely extracted from input LF. Finally, considering that self-supervised methods need to retrain from the beginning toward each test image, based on the proposed “part-to-part” strategy, an efficient end-to-end network is designed to extract these cross-view features for superior SASR performance with less training time. Experiment results demonstrate that our method outperforms other state-of-the-art ESLB methods on both large and small domain gap cases. Compared with the only SSLB method (LFZSSR), our approach achieves better performance with 524 times less training time.
Hao Sheng 0001, Da Yang 0001, Ruixuan Cong, Zhenglong Cui, Rongshan Chen
IEEE Trans. Circuits Syst. Video Technol.4
2022 UrbanLF: A Comprehensive Light Field Dataset for Semantic Segmentation of Urban Scenes
abstract
As one of the fundamental technologies for scene understanding, semantic segmentation has been widely explored in the last few years. Light field cameras encode the geometric information by simultaneously recording the spatial information and angular information of light rays, which provides us with a new way to solve this issue. In this paper, we propose a high-quality and challenging urban scene dataset, containing 1074 samples composed of real-world and synthetic light field images as well as pixel-wise annotations for 14 semantic classes. To the best of our knowledge, it is the largest and the most diverse light field dataset for semantic segmentation. We further design two new semantic segmentation baselines tailored for light field and compare them with state-of-the-art RGB, video and RGB-D-based methods using the proposed dataset. The outperforming results of our baselines demonstrate the advantages of the geometric information in light field for this task. We also provide evaluations of super-resolution and depth estimation methods, showing that the proposed dataset presents new challenges and supports detailed comparisons among different methods. We expect this work inspires new research direction and stimulates scientific progress in related fields. The complete dataset is available athttps://github.com/HAWKEYE-Group/UrbanLF.
Hao Sheng 0001, Ruixuan Cong, Da Yang 0001, Rongshan Chen, Zhenglong Cui
IEEE Trans. Circuits Syst. Video Technol.2