EDBT 2026 Demo / reviewers in the wild / expert
Shuai Guo 0002
dblp:07/272-2
· DBLP profile ↗
12ranked-venue papers
7as first author
9since 2021 · last 2024
0000-0001-9102-6545ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 8 · 5 first-author · 8 since 2021Artificial intelligence and machine learning · 4 · 3 first-author · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Depth-Guided Robust and Fast Point Cloud Fusion NeRF for Sparse Input ViewsabstractNovel-view synthesis with sparse input views is important for real-world applications like AR/VR and autonomous driving. Recent methods have integrated depth information into NeRFs for sparse input synthesis, leveraging depth prior for geometric and spatial understanding. However, most existing works tend to overlook inaccuracies within depth maps and have low time efficiency. To address these issues, we propose a depth-guided robust and fast point cloud fusion NeRF for sparse inputs. We perceive radiance fields as an explicit voxel grid of features. A point cloud is constructed for each input view, characterized within the voxel grid using matrices and vectors. We accumulate the point cloud of each input view to construct the fused point cloud of the entire scene. Each voxel determines its density and appearance by referring to the point cloud of the entire scene. Through point cloud fusion and voxel grid fine-tuning, inaccuracies in depth values are refined or substituted by those from other views. Moreover, our method can achieve faster reconstruction and greater compactness through effective vector-matrix decomposition. Experimental results underline the superior performance and time efficiency of our approach compared to state-of-the-art baselines. Shuai Guo 0002, Qiuwen Wang, Yijie Gao, Rong Xie 0004, Li Song 0001 |
AAAI | 1 |
| 2024 | A New People-Object Interaction Dataset and NVS BenchmarksabstractRecently, NVS in human-object interaction scenes has received increasing attention. Existing human-object interaction datasets mainly consist of static data with limited views, offering only RGB images or videos, mostly containing interactions between a single person and objects. Moreover, these datasets exhibit complexities in lighting environments, poor synchronization, and low resolution, hindering high-quality human-object interaction studies. In this paper, we introduce a new people-object interaction dataset that comprises 38 series of 30-view multi-person or single-person RGB-D video sequences, accompanied by camera parameters, foreground masks, SMPL models, some point clouds, and mesh files. Video sequences are captured by 30 Kinect Azures, uniformly surrounding the scene, each in 4 K resolution 25 FPS, and lasting for 1~19 seconds. Meanwhile, we evaluate some SOTA NVS models on our dataset to establish the NVS benchmarks. We hope our work can inspire further research in human-object interaction. Shuai Guo 0002, Houqiang Zhong, Qiuwen Wang, Yijie Gao, Jiajing Yuan, Rong Xie 0004, Li Song 0001 |
ICIP | 1 |
| 2024 | Depth-Guided Robust Point Cloud Fusion NeRF for Sparse Input ViewsabstractNovel-view synthesis with sparse input views is important for practical applications such as AR/VR and autonomous driving. Many works in this field have already integrated depth information into NeRF, utilizing depth priors for assistance in geometric and spatial understanding. However, most existing work tends to either overlook the inaccuracies in depth maps or only handle them roughly, limiting the effectiveness of the synthesis. To address this issue, we propose a depth-guided robust point cloud fusion NeRF for sparse input synthesis. We first construct a point cloud for each input view, with a novel point cloud representation based on learnable matrices and vectors. Then, through an additional lightweight scene fusion network, we fuse the point clouds from each input view to build a point cloud of the entire scene. By optimizing the point cloud representation and scene fusion network, inaccuracies in the depth map can be adjusted and refined, thereby achieving a more precise perception of the overall scene. Each voxel in the scene is determined by referencing the fused point cloud to establish its density and appearance. Experimental results demonstrate that our method outperforms state-of-the-art baselines. Shuai Guo 0002, Qiuwen Wang, Yijie Gao, Rong Xie 0004, Lin Li 0062, Li Song 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2024 | Real-Time Free Viewpoint Video Synthesis System Based on DIBR and a Depth Estimation NetworkabstractDepth image-based rendering (DIBR) view synthesis is the most widely employed method in real-time FVV research. Despite recent progress, most DIBR-based FVV synthesis approaches are not sufficiently simple and effective in filling holes and artifacts. Additionally, they use RGB-D cameras, which are difficult to widely adopt or take considerable time to estimate high-quality depth images. This paper introduces a real-time FVV synthesis system based on DIBR and a depth estimation network. This system includes a 12-view synchronous camera system, a new multistage depth estimation network, a new GPU-accelerated DIBR algorithm, and a virtual view parameter generation method. This system provides the first real-time FVV solution for background-fixed fields based on DIBR and a depth estimation network. It can infer depth images for all camera views and synthesize any virtual view along the horizontal circular arc of the camera rig in real time. To our knowledge, we are the first to introduce background models and foreground masks and a refined multistage structure to address real-time high-quality depth estimation and DIBR FVV synthesis. We also build a high-quality multiview RGB-D synchronous dataset that has promising DIBR FVV synthesis performance to train and evaluate our system. The experimental results demonstrate the real-time and better performance of the proposed system. Shuai Guo 0002, Jingchuan Hu, Kai Zhou 0016, Jionghao Wang, Li Song 0001, Rong Xie 0004, Wenjun Zhang 0001 |
IEEE Trans. Multim. | 1 |
| 2023 | NeRF-SDP: Efficient Generalizable Neural Radiance Field with Scene Depth PerceptionabstractIn recent years, neural radiance fields have exhibited impressive performance in novel view synthesis. However, exploiting complex network structures to achieve generalizable NeRF usually results in inefficient rendering. Existing methods for accelerating rendering directly employ simpler inference networks or fewer sampling points, leading to unsatisfactory synthesis quality. To address the challenge of balancing rendering speed and quality in generalizable NeRF, we propose a novel framework, NeRF-SDP, which achieves both efficiency and high fidelity by introducing scene depth perception. We incorporate more scene information into the radiance field by using our proposed geometry feature extraction and depth-encoded ray transformer to improve the model’s inference capabilities with sparse points. With the aid of scene depth perception, NeRF-SDP can better understand the scene’s structure, thus better reconstructing the objects’ edges with significantly fewer artifacts. Experimental results demonstrate that NeRF-SDP achieves comparable synthesis quality to state-of-the-art methods while significantly improving rendering efficiency. Furthermore, ablation studies confirm that the depth-encoded ray transformer enhances the model’s robustness to varying numbers of sampling points. Qiuwen Wang, Shuai Guo 0002, Haoning Wu 0002, Rong Xie 0004, Li Song 0001, Wenjun Zhang 0001 |
MMAsia | 2 |
| 2022 | A Multi-User Oriented Live Free-Viewpoint Video Streaming System Based on View InterpolationabstractAs an important application form of immersive multimedia services, free-viewpoint video (FVV) enables users with great immersive experience by strong interaction. However, the computational complexity of virtual view synthesis algorithms poses a significant challenge to the real-time performance of an FVV system. Furthermore, the individuality of user interaction makes it difficult to serve multiple users simultaneously for a system with conventional architecture. In this paper, we novelly introduce a CNN-based view interpolation algorithm to synthesis dense virtual views in real time. Based on this, we also build an end-to-end live free-viewpoint system with a multi-user oriented streaming strategy. Our system can utilize a single edge server to serve multiple users at the same time without having to bring a large view synthesis load on the client side. We analyze the whole system and show that our approaches give the user a pleasant immersive experience, in terms of both visual quality and latency. Jingchuan Hu, Shuai Guo 0002, Kai Zhou 0016, Jun Xu 0040, Li Song 0001 |
ICME | 2 |
| 2022 | A new free viewpoint video dataset and DIBR benchmarkabstractFree viewpoint video (FVV) has drawn great attention in recent years, which provides viewers with strong interactive and immersive experience. Despite the developments made, further progress of FVV research is limited by existing datasets that mostly have too few number of camera views, or static scenes. To overcome the limitations, in this paper, we present a new dynamic RGB-D video dataset with up to 12 views. Our dataset consists of 13 groups of dynamic video sequences that are taken at the same scene, and a group of video sequences of the empty scene. Each group has 12 HD video sequences taken by synchronized cameras and 12 correspondingly estimated depth video sequences. Moreover, we also introduce a FVV synthesis benchmark on the basis of depth image based rendering (DIBR) to help researchers validate their data-driven methods. We hope our work will inspire more FVV synthesis methods with enhanced robustness, improved performance and deeper understanding. Shuai Guo 0002, Kai Zhou 0016, Jingchuan Hu, Jionghao Wang, Jun Xu 0040, Li Song 0001 |
MMSys | 1 |
| 2022 | RGBD-based Real-time Volumetric Reconstruction System: Architecture Design and ImplementationabstractWith the increasing popularity of commercial depth cameras, 3D reconstruction of dynamic scenes has aroused widespread interest. Although many novel 3D applications have been unlocked, real-time performance is still a big problem. In this paper, a low-cost, real-time system: LiveRecon3D, is presented, with multiple RGB-D cameras connected to one single computer. The goal of the system is to provide an interactive frame rate for 3D content capture and rendering at a reduced cost. In the proposed system, we adopt a scalable volume structure and employ ray casting technique to extract the surface of 3D content. Based on a pipeline design, all the modules in the system run in parallel and are designed to minimize the latency to achieve an interactive frame rate of 30 FPS. At last, experimental results corresponding to implementation with three Kinect v2 cameras are presented to verify the system's effectiveness in terms of visual quality and real-time performance. Kai Zhou 0016, Shuai Guo 0002, Jingchuan Hu, Jionghao Wang, Qiuwen Wang, Li Song 0001 |
VCIP | 2 |
| 2022 | Multiview nonlinear discriminant structure learning for emotion recognition
Shuai Guo 0002, Li Song 0001, Rong Xie 0004, Lin Li 0062, Shenglan Liu 0001 |
Knowl. Based Syst. | 1 |
| 2020 | An incrementally cascaded broad learning framework to facial landmark tracking
Caifeng Liu, Lin Feng 0001, Shuai Guo 0002, Huibing Wang, Shenglan Liu 0001, Hong Qiao |
Neurocomputing | 3 |
| 2020 | Multi-view laplacian eigenmaps based on bag-of-neighbors for RGB-D human emotion recognition
Shenglan Liu 0001, Shuai Guo 0002, Wei Wang 0036, Hong Qiao, Wenbo Luo |
Inf. Sci. | 2 |
| 2019 | Multi-view laplacian least squares for human emotion recognition
Shuai Guo 0002, Lin Feng 0001, Zhanbo Feng, Yi-Hao Li, Shenglan Liu 0001, Hong Qiao |
Neurocomputing | 1 |