VLDB 2026 Research / reviewers in the wild / expert
Tun Wang
dblp:178/9530
· DBLP profile ↗
10ranked-venue papers
4as first author
10since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021Security and privacy · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Difference-guided full-view volume for light field depth estimation
Tun Wang, Hao Sheng 0001, Ruixuan Cong, Da Yang 0001, Zhenglong Cui, Guanqun Su |
Expert Syst. Appl. | 1 |
| 2026 | Gradient-Guided Density Redistribution Network for Light Field Full-View Depth EstimationabstractLight field (LF) full-view depth estimation aims to recover dense and coherent depth maps for all sub-aperture views, which is crucial for applications such as 3D reconstruction, LF editing and virtual reality. However, directly extending center-view volume-based methods to the full-view is computationally infeasible, as it requires constructing a separate cost volume for each view. Besides, existing full-view propagation-based approaches, while more efficient, frequently suffer from edge fattening and cross-view inconsistencies in the presence of occlusions. In this paper, we propose a gradient-guided density redistribution network (GDRNet), a novel end-to-end framework that efficiently generates full-view depth maps by constructing a single plane-density volume and a multi-plane depth image, which are then propagated to all angular views. To resolve ambiguous estimates at occlusion edges, we perform a direction-aware gradient-guided density redistribution only inside a dilated edge narrow band. For each center pixel in edge regions, a guidance gradient is derived from the initial depth map to determine the normal and tangent directions. Then, density in edge fattening regions can be redistributed via sampling along the normal direction, while similarity along the tangent direction can fill bad pixels with inconsistencies. Furthermore, an adaptive edge extraction module with four directional learnable Sobel kernels is designed to jointly exploit spatial and angular gradients, enabling robust detection and localizing the refinement band. Extensive experiments on synthetic and real-world LF datasets demonstrate that GDRNet achieves state-of-the-art accuracy and edge sharpness in both quantitative and qualitative evaluations, while maintaining computational efficiency compared to full-view methods. Tun Wang, Zhenglong Cui, Ruixuan Cong, Da Yang 0001, Mingyuan Zhao 0001, Hao Sheng 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2026 | UCGR: Closing the Discretization Gap in Light Field Depth Estimation via Unified Continuous Geometry RepresentationabstractLight field (LF) cameras encode dense spatial-angular information for depth estimation, critical for applications such as 3D reconstruction, refocusing, and virtual reality. However, current deep learning methods for LF depth estimation still face significant challenges due to the discretization gap between the continuous geometry of real scenes and the discrete sampling of digital images, limiting their effectiveness in high-precision application. This gap manifests in two complementary forms: spatial discretization leads to structural ambiguities and information loss, while depth discretization introduces inaccuracies due to fixed, discrete depth sampling. To address these challenges, we propose Unified Continuous Geometry Representation (UCGR), a unified representation that models scene geometry as a continuous field over image coordinates and depth. UCGR treats spatial and depth discretization as two facets of the same problem and realized by two complementary operators: (1) Adaptive Plane Sampling Operator, which learns edge-aware planar priors to preserve geometric details and mitigate spatial discretization. (2) Contextual Depth Correction Operator, which utilizes contextual information for depth correction, ensuring continuous depth estimation and suppressing artifacts. Building on UCGR, we propose a Continuous Geometry Network (CGNet) that collaboratively optimizes both spatial and depth discretization for accurate and consistent LF depth estimation. Extensive experiments on synthetic and real-world LF datasets demonstrate that CGNet achieves state-of-the-art performance, significantly outperforming existing LF depth estimation methods in terms of both accuracy and robustness. Zexin Sun, Tun Wang, Rongshan Chen, Ruixuan Cong, Wei Ke 0001, Hao Sheng 0001 |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2025 | Depth State Space Model for Light Field Depth Estimation via Text-Similar Representation
Zexin Sun, Tun Wang, Da Yang 0001, Zhenglong Cui, Rongshan Chen, Ying Li 0122, Guanqun Su, Hao Sheng 0001 |
KSEM (1) | 2 |
| 2025 | Multiplane depth image for view-consistent light field depth estimation
Tun Wang, Hao Sheng 0001, Rongshan Chen, Ruixuan Cong, Mingyuan Zhao 0001, Da Yang 0001 |
Knowl. Based Syst. | 1 |
| 2025 | A GPU-Enabled Framework for Light Field Efficient Compression and Real-Time RenderingabstractReal-time rendering offers instantaneous visual feedback, making it crucial for mixed-reality applications. The light field captures both light intensity and direction in a 3D environment, serving as a data-rich medium to enhance mixed-reality experiences. However, two major challenges remain: 1) current light field rendering techniques are unsuitable for real-time computation, and 2) existing real-time methods cannot efficiently process high-dimensional light field data on GPU platforms. To overcome these challenges, we propose an framework utilizing a compact neural representation of light field data, implemented on a GPU platform for real-time rendering. This framework provides both compact storage and high-fidelity real-time computation. Specifically, we introduce a ray global alignment strategy to simplify the framework and improve practicality. This strategy enables the learning of an optimal embedding for all local rays in a globally consistent way, removing the need for camera pose calculations. To achieve effective compression, the neural light field is employed to map each embedded ray to its corresponding color. To enable real-time rendering, we design a novel super-resolution network to enhance rendering speed. Extensive experiments demonstrate that our framework significantly enhances compression efficiency and real-time rendering performance, achieving nearly 50$\mathbf{\times}$compression ratio and 100 FPS rendering. Mingyuan Zhao 0001, Hao Sheng 0001, Rongshan Chen, Ruixuan Cong, Tun Wang, Zhenglong Cui, Da Yang 0001, Shuai Wang 0027, Wei Ke 0001 |
IEEE Trans. Computers | 5 |
| 2025 | Towards Depth-Continuous Scene Representation With a Displacement Field for Robust Light Field Depth EstimationabstractLight field (LF) captures both spatial and angular information of scenes, enabling accurate depth estimation. However, previous deep learning methods have typically model surface depth only, while ignoring the continuous nature of depth in 3D scenes. In this paper, we use displacement field (DF) to describe this continuous property, and propose a novel depth-continuous scene representation for robust LF depth estimation. Experiments demonstrate that our representation enables the network to generate highly detailed depth maps with fewer parameters and faster speed. Specifically, inspired by signed distance field in 3D object description, we aim to exploit the intrinsic depth-continuous property of 3D scenes using DF, and define a novel depth-continuous scene representation. Then, we introduce a simple yet general learning framework for depth-continuous scene embedding, and the proposed network, DepthDF, achieves state-of-the-art performance on both synthetic and real-world LF datasets, ranking 1st on the HCI 4D Light Field benchmark. Furthermore, previous LF depth estimation methods can also be seamlessly integrated into this framework. Finally, we extend this framework beyond LF depth estimation to various tasks, including multi-view stereo depth inference, LF super-resolution, and LF salient object detection. Experiments demonstrate improved performance when the continuous scene representation is applied, suggesting that our framework can potentially bring insights to more fields. Rongshan Chen, Hao Sheng 0001, Da Yang 0001, Ruixuan Cong, Zhenglong Cui, Tun Wang, Mingyuan Zhao 0001 |
IEEE Trans. Multim. | 7 |
| 2024 | Light field depth estimation: A comprehensive survey from principles to futureabstractLight field (LF) depth estimation is an important research direction in the area of computer vision and computational photography, which aims to infer the depth information of different objects in three-dimensional scenes by capturing LF data. Given this new era of significance, this article introduces a survey of the key concepts, methods, novel applications, and future trends in this area. We summarize the LF depth estimation methods, which are usually based on the interaction of radiance from rays in all directions of the LF data, such as epipolar-plane, multi-view geometry, focal stack, and deep learning. We analyze the many challenges facing each of these approaches, including complex algorithms, large amounts of computation, and speed requirements. In addition, this survey summarizes most of the currently available methods, conducts some comparative experiments, discusses the results, and investigates the novel directions in LF depth estimation. Tun Wang, Hao Sheng 0001, Rongshan Chen, Da Yang 0001, Zhenglong Cui, Ruixuan Cong, Mingyuan Zhao 0001 |
High Confid. Comput. | 1 |
| 2024 | A survey for light field super-resolutionabstractCompared to 2D imaging data, the 4D light field (LF) data retains richer scene’s structure information, which can significantly improve the computer’s perception capability, including depth estimation, semantic segmentation, and LF rendering. However, there is a contradiction between spatial and angular resolution during the LF image acquisition period. To overcome the above problem, researchers have gradually focused on the light field super-resolution (LFSR). In the traditional solutions, researchers achieved the LFSR based on various optimization frameworks, such as Bayesian and Gaussian models. Deep learning-based methods are more popular than conventional methods because they have better performance and more robust generalization capabilities. In this paper, the present approach can mainly divided into conventional methods and deep learning-based methods. We discuss these two branches in light field spatial super-resolution (LFSSR), light field angular super-resolution (LFASR), and light field spatial and angular super-resolution (LFSASR) , respectively. Subsequently, this paper also introduces the primary public datasets and analyzes the performance of the prevalent approaches on these datasets. Finally, we discuss the potential innovations of the LFSR to propose the progress of our research field. Mingyuan Zhao 0001, Hao Sheng 0001, Da Yang 0001, Ruixuan Cong, Zhenglong Cui, Rongshan Chen, Tun Wang, Shuai Wang 0027 |
High Confid. Comput. | 8 |
| 2024 | Multimodal Perception Integrating Point Cloud and Light Field for Ship Autonomous DrivingabstractRobust scene perception is an essential prerequisite to ensure the reliability in ship autonomous driving. However, it is a challenging task in inland river because of the complicated and changeable environment as well as high-density ships in narrow waterway. As one of the primary technologies, obstacle trajectory locating and tracking has been widely explored in recent years. Current approaches strictly rely on lidar as only depth awareness sensor and the limited measurement range severely restricts them for distant object identification. On this account, we creatively propose a point cloud-light field fusion perception framework in this paper for the first time. Specifically, in detection stage, the former undertakes precise close object perception and the latter completes distant object locating through light field stereo matching. In tracking stage, a novel four-phase data association that combines multiple attributes from position, point cloud and image domains is utilized for accurate object matching across frames. To validate the effectiveness of our multimodal perception strategy, we implement an acquisition system consisting of two lidars and four sets of simplified light field cameras on a ship to conduct actual testing. Extensive experimental results show that the proposed framework achieves superior 3D object locating and tracking performance, far surpassing the state-of-the-art methods in terms of accuracy and real-time. Ruixuan Cong, Hao Sheng 0001, Mingyuan Zhao 0001, Dazhi Yang 0003, Tun Wang, Rongshan Chen |
IEEE Trans. Intell. Transp. Syst. | 5 |