Kai Zhou 0016

dblp:432/0541 · DBLP profile ↗
← Back
4ranked-venue papers
1as first author
4since 2021 · last 2024
0009-0009-7075-5656ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 4 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer graphics and multimedia
1 paper
Rendering · 56% Virtual and augmented reality · 44%
Artificial intelligence
1 paper
3D vision · 100%

Topics — the 4 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Rendering › image-based rendering
depth-image-based rendering
0.812024
Real-Time Free Viewpoint Video Synthesis System Based on DIBR and a Depth Estimation Network · IEEE Trans. Multim. 2024
Virtual and augmented reality › immersive video
free-viewpoint video
0.812024
Real-Time Free Viewpoint Video Synthesis System Based on DIBR and a Depth Estimation Network · IEEE Trans. Multim. 2024
Computer vision › 3D vision
depth estimation
0.212024
Real-Time Free Viewpoint Video Synthesis System Based on DIBR and a Depth Estimation Network · IEEE Trans. Multim. 2024
Rendering
novel view synthesis
0.212024
Real-Time Free Viewpoint Video Synthesis System Based on DIBR and a Depth Estimation Network · IEEE Trans. Multim. 2024

Methods — techniques the papers use, named apart from their topics

depth estimation network · 1.5GPU-accelerated rendering · 1.5
YearPublicationVenuePosition
2024 Real-Time Free Viewpoint Video Synthesis System Based on DIBR and a Depth Estimation Network
abstract
Depth image-based rendering (DIBR) view synthesis is the most widely employed method in real-time FVV research. Despite recent progress, most DIBR-based FVV synthesis approaches are not sufficiently simple and effective in filling holes and artifacts. Additionally, they use RGB-D cameras, which are difficult to widely adopt or take considerable time to estimate high-quality depth images. This paper introduces a real-time FVV synthesis system based on DIBR and a depth estimation network. This system includes a 12-view synchronous camera system, a new multistage depth estimation network, a new GPU-accelerated DIBR algorithm, and a virtual view parameter generation method. This system provides the first real-time FVV solution for background-fixed fields based on DIBR and a depth estimation network. It can infer depth images for all camera views and synthesize any virtual view along the horizontal circular arc of the camera rig in real time. To our knowledge, we are the first to introduce background models and foreground masks and a refined multistage structure to address real-time high-quality depth estimation and DIBR FVV synthesis. We also build a high-quality multiview RGB-D synchronous dataset that has promising DIBR FVV synthesis performance to train and evaluate our system. The experimental results demonstrate the real-time and better performance of the proposed system.
Shuai Guo 0002, Jingchuan Hu, Kai Zhou 0016, Jionghao Wang, Li Song 0001, Rong Xie 0004, Wenjun Zhang 0001
IEEE Trans. Multim.3
2022 A Multi-User Oriented Live Free-Viewpoint Video Streaming System Based on View Interpolation
abstract
As an important application form of immersive multimedia services, free-viewpoint video (FVV) enables users with great immersive experience by strong interaction. However, the computational complexity of virtual view synthesis algorithms poses a significant challenge to the real-time performance of an FVV system. Furthermore, the individuality of user interaction makes it difficult to serve multiple users simultaneously for a system with conventional architecture. In this paper, we novelly introduce a CNN-based view interpolation algorithm to synthesis dense virtual views in real time. Based on this, we also build an end-to-end live free-viewpoint system with a multi-user oriented streaming strategy. Our system can utilize a single edge server to serve multiple users at the same time without having to bring a large view synthesis load on the client side. We analyze the whole system and show that our approaches give the user a pleasant immersive experience, in terms of both visual quality and latency.
Jingchuan Hu, Shuai Guo 0002, Kai Zhou 0016, Jun Xu 0040, Li Song 0001
ICME4
2022 A new free viewpoint video dataset and DIBR benchmark
abstract
Free viewpoint video (FVV) has drawn great attention in recent years, which provides viewers with strong interactive and immersive experience. Despite the developments made, further progress of FVV research is limited by existing datasets that mostly have too few number of camera views, or static scenes. To overcome the limitations, in this paper, we present a new dynamic RGB-D video dataset with up to 12 views. Our dataset consists of 13 groups of dynamic video sequences that are taken at the same scene, and a group of video sequences of the empty scene. Each group has 12 HD video sequences taken by synchronized cameras and 12 correspondingly estimated depth video sequences. Moreover, we also introduce a FVV synthesis benchmark on the basis of depth image based rendering (DIBR) to help researchers validate their data-driven methods. We hope our work will inspire more FVV synthesis methods with enhanced robustness, improved performance and deeper understanding.
Shuai Guo 0002, Kai Zhou 0016, Jingchuan Hu, Jionghao Wang, Jun Xu 0040, Li Song 0001
MMSys2
2022 RGBD-based Real-time Volumetric Reconstruction System: Architecture Design and Implementation
abstract
With the increasing popularity of commercial depth cameras, 3D reconstruction of dynamic scenes has aroused widespread interest. Although many novel 3D applications have been unlocked, real-time performance is still a big problem. In this paper, a low-cost, real-time system: LiveRecon3D, is presented, with multiple RGB-D cameras connected to one single computer. The goal of the system is to provide an interactive frame rate for 3D content capture and rendering at a reduced cost. In the proposed system, we adopt a scalable volume structure and employ ray casting technique to extract the surface of 3D content. Based on a pipeline design, all the modules in the system run in parallel and are designed to minimize the latency to achieve an interactive frame rate of 30 FPS. At last, experimental results corresponding to implementation with three Kinect v2 cameras are presented to verify the system's effectiveness in terms of visual quality and real-time performance.
Kai Zhou 0016, Shuai Guo 0002, Jingchuan Hu, Jionghao Wang, Qiuwen Wang, Li Song 0001
VCIP1