Zhihui Ke

dblp:276/4518 · DBLP profile ↗
← Back
11ranked-venue papers
3as first author
11since 2021 · last 2026
0009-0002-4042-7044ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 6 since 2021Artificial intelligence and machine learning · 5 · 1 first-author · 5 since 2021Computer networks · 3 · 1 first-author · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 StreamSTGS: Streaming Spatial and Temporal Gaussian Grids for Real-Time Free-Viewpoint Video
abstract
Streaming free-viewpoint video (FVV) in real-time still faces significant challenges, particularly in training, rendering, and transmission efficiency. Harnessing superior performance of 3D Gaussian Splatting (3DGS), recent 3DGS-based FVV methods have achieved notable breakthroughs in both training and rendering. However, the storage requirements of these methods can reach up to 10MB per frame, making stream FVV in real-time impossible. To address this problem, we propose a novel FVV representation, dubbed StreamSTGS, designed for real-time streaming. StreamSTGS represents a dynamic scene using canonical 3D Gaussians, temporal features, and a deformation field. For high compression efficiency, we encode canonical Gaussian attributes as 2D images and temporal features as a video. This design not only enables real-time streaming, but also inherently supports adaptive bitrate control based on network condition without any extra training. Moreover, we propose a sliding window scheme to aggregate adjacent temporal features to learn local motions, and then introduce a transformer-guided auxiliary training module to learn global motions. On diverse FVV benchmarks, StreamSTGS demonstrates competitive performance on all metrics compared to state-of-the-art methods. Notably, StreamSTGS increases the PSNR by an average of 1dB while reducing the average frame size to just 170KB.
Zhihui Ke, Yvyang Liu, Xiaobo Zhou 0003, Tie Qiu 0001
AAAI1
2025 4D-Editor: Interactive Object-Level Editing in Dynamic Neural Radiance Fields via Semantic Distillation
abstract
This paper targets interactive object-level editing (e.g., deletion, recoloring, transformation, composition) in dynamic scenes. Recently, some methods aiming for flexible editing static scenes represented by neural radiance field (NeRF) have shown impressive synthesis quality, while similar capabilities in time-variant dynamic scenes remain limited. To solve this problem, we propose 4D-Editor, an interactive semantic-driven editing framework, allowing editing multiple objects in a dynamic NeRF with user strokes on a single frame. Specifically, we extend the original dynamic NeRF by incorporating Hybrid Semantic Feature Distillation to maintain spatial-temporal consistency after editing. In addition, a Recursive Selection Refinement module is presented to significantly boost object segmentation accuracy within a dynamic NeRF to aid the editing process. Moreover, we develop Multi-view Reprojection Inpainting to fill holes caused by incomplete scene capture after editing. Extensive quantitative and qualitative experiments on real application scenarios demonstrate that 4D-Editor achieves photo-realistic editing on dynamic NeRFs. Project page: https://patrickddj.github.ioI/4D-Editor
Dadong Jiang, Zhihui Ke, Xiaobo Zhou 0003, Tie Qiu 0001, Xidong Shi
3DV2
2025 FlexiTex: Enhancing Texture Generation via Visual Guidance
abstract
Recent texture generation methods achieve impressive results due to the powerful generative prior they leverage from large-scale text-to-image diffusion models. However, abstract textual prompts are limited in providing global textural or shape information, which results in the texture generation methods producing blurry or inconsistent patterns. To tackle this, we present FlexiTex, embedding rich information via visual guidance to generate a high-quality texture. The core of FlexiTex is the Visual Guidance Enhancement module, which incorporates more specific information from visual guidance to reduce ambiguity in the text prompt and preserve high-frequency details. To further enhance the visual guidance, we introduce a Direction-Aware Adaptation module that automatically designs direction prompts based on different camera poses, avoiding the Janus problem and maintaining semantically global consistency. Benefiting from the visual guidance, FlexiTex produces quantitatively and qualitatively sound results, demonstrating its potential to advance texture generation for real-world applications.
Dadong Jiang, Xianghui Yang, Zibo Zhao 0001, Zeqiang Lai, Shaoxiong Yang, Chunchao Guo, Xiaobo Zhou 0003, Zhihui Ke
AAAI10
2025 Communication-Efficient Multi-Vehicle Collaborative Semantic Segmentation via Sparse 3D Gaussian Sharing
Tianyu Hong, Xiaobo Zhou 0003, Wenkai Hu, Qi Xie 0003, Zhihui Ke, Tie Qiu 0001
ICCV5
2025 Timeformer: Capturing Temporal Relationships of Deformable 3D Gaussians for Robust Reconstruction
abstract
Dynamic scene reconstruction is a long-term challenge in 3D vision. Recent methods extend 3D Gaussian Splatting to dynamic scenes via additional deformation fields and apply explicit constraints like motion flow to guide the deformation. However, they learn motion changes from individual timestamps independently, making it challenging to reconstruct complex scenes, particularly when dealing with violent movement, extreme-shaped geometries, or reflective surfaces. To address the above issue, we design a plug-and-play module called TimeFormer to enable existing deformable 3D Gaussians reconstruction methods with the ability to implicitly model motion patterns from a learning perspective. Specifically, TimeFormer includes a Cross-Temporal Transformer Encoder, which adaptively learns the temporal relationships of deformable 3D Gaussians. Furthermore, we propose a two-stream optimization strategy that transfers the motion knowledge learned from TimeFormer to the base stream during the training phase. This allows us to remove TimeFormer during inference, thereby preserving the original rendering speed. Extensive experiments in the multi-view and monocular dynamic scenes validate qualitative and quantitative improvement brought by TimeFormer. Project Page: https://patrickddj.github.io/TimeFormer/
Dadong Jiang, Zhi Hou, Zhihui Ke, Xianghui Yang, Xiaobo Zhou 0003, Tie Qiu 0001
ICCV3
2025 EdgeGaussian: Real-time Free-Viewpoint Video for Mobile VR via Edge-Client Collaborative Neural Rendering
abstract
Free-Viewpoint Videos (FVVs) enable immersive viewing of a scene from any position and angle using virtual reality (VR) head-mounted displays (HMDs), thus have great potential in various applications such as telepresence, gaming, and education. Recently, 3D Gaussian Splatting (3DGS) has emerged as a promising method for FVV construction due to its superior reconstruction quality. However, its real-time rendering on untethered HMDs remains challenging due to high computational demands. To address this challenge, we propose EdgeGaussian, a novel edge-client collaborative framework for real-time FVV rendering. Our approach employs decomposed static-dynamic 4D Gaussian splatting (SD-4DGS) to separately reconstruct static and dynamic components of a scene. We further introduce a hybrid mesh-4DGS neural representation, where static components are modeled as textured meshes for local rendering, while dynamic components are offloaded to edge servers as 4DGS. This decomposition significantly reduces the computational burden on the client device while maintaining high rendering quality. Our testbed experiments demonstrate that Edge-Gaussian achieves up to 128 FPS, outperforming state-of-the-art local rendering methods by 4x and edge rendering methods by 5x.
Zhihui Ke, Xiaobo Zhou 0003, Zhizhuo Pang, Tie Qiu 0001
MobiCom1
2024 DS-NeRV: Implicit Neural Video Representation with Decomposed Static and Dynamic Codes
abstract
Implicit neural representations for video (NeRV) have recently become a novel way for high-quality video representation. However, existing works employ a single network to represent the entire video, which implicitly con-fuse static and dynamic information. This leads to an inability to effectively compress the redundant static information and lack the explicitly modeling of global temporal-coherent dynamic details. To solve above problems, we propose DS-NeRV, which decomposes videos into sparse learnable static codes and dynamic codes without the need for explicit optical flow or residual supervision. By setting different sampling rates for two codes and applying weighted sum and interpolation sampling methods, DS-NeRV efficiently utilizes redundant static information while maintaining high-frequency details. Additionally, we design a cross-channel attention-based (CCA) fusion module to efficiently fuse these two codes for frame decoding. Our approach achieves a high quality reconstruction of 31.2 PSNR with only 0.35M parameters thanks to separate static and dynamic codes representation and outperforms existing NeRV methods in many downstream tasks. Our project website is at https://haoyan14.github.io/DS-NeRV/.
Zhihui Ke, Xiaobo Zhou 0003, Tie Qiu 0001, Xidong Shi, Dadong Jiang
CVPR2
2024 TP-BFT: A Faster Asynchronous BFT Consensus with Parallel Structure
Shunliang Ye, Fuan Xiao, Zhihui Ke, Guoyu Yang, Huawei Ma
ICA3PP (4)4
2024 Recommendation-Driven Multi-Cell Cooperative Caching: A Multi-Agent Reinforcement Learning Approach
abstract
In 5 G small cell networks, edge caching is a key technique to alleviate the backhaul burden by caching user desired contents at network edges such as small base stations (SBSs). However, due to storage space limitation and diverse user preference patterns, a single SBS is unable to cache all the user desired contents and thus leading to low caching efficiency. In this paper, we propose a recommendation-driven multi-cell cooperative caching strategy to improve the caching efficiency. The idea is to aggregate the storage spaces of multiple SBSs into a large shared resource pool, and guide users to access cached contents by content recommendation. First, we formulate the joint cooperative caching and recommendation problem as a multi-agent multi-armed bandit (MAMAB) problem with the aim of minimizing the average download latency. Then, we propose a multi-agent reinforcement learning (MARL)-based algorithm, MARL-JCR, to solve the problem in a fully distributed manner with limited information exchange among the agents. We also develop a modified combinatorial upper confidence bound algorithm to reduce each agent's decision space to reduce computational complexity. The experiment results evaluated on theMovieLensdataset show MARL-JCR decreases the average download latency by up to 60% as compared with the state-of-the-art solutions.
Xiaobo Zhou 0003, Zhihui Ke, Tie Qiu 0001
IEEE Trans. Mob. Comput.2
2023 CollabVr: Reprojection-Based Edge-Client Collaborative Rendering for Real-Time High-Quality Mobile Virtual Reality
abstract
Collaborative mobile virtual reality (VR) has recently emerged as a promising solution to provide an immersive user experience with low motion-to-photo (MTP) latency. The rendering tasks are usually divided into background and foreground ones, which are executed in the edge server and head-mounted display (HMD), respectively. Assuming that the background images are static, they can be reused for temporal redundancy reduction in transmission. However, in high dynamic high-quality scenes, background images are continuously changing, making the temporal reuse strategy ineffective, leading to high MTP latency and hence motion sickness. In this paper, we propose CollabVR, a reprojection-based edge-client collaborative rendering approach for real-time high-quality mobile VR in high dynamic scenes. The key idea is to reduce the spatial redundancy in transmission by exploiting the high similarity between the left view and right view. With CollabVR, only one view is rendered and encoded in the edge server and then be transmitted to, and decoded at the HMD. The another view is reprojected by utilizing depth image-based rendering (DIBR) in the HMD, thereby greatly reducing the MTP latency even in high dynamic scenes. Furthermore, we propose a foveated-based multi-level patch subdivision strategy to achieve real-time reprojection in the resourced-limited HMD. A parallel streaming strategy is also proposed to fill holes that exist in the reprojected image. Experiments we conducted using Commercial Off-The-Shelf (COTS) devices indicate that CollabVR can reduce the average MTP latency by up to 36% compared to the baseline methods.
Zhihui Ke, Xiaobo Zhou 0003, Dadong Jiang, Tie Qiu 0001
RTSS1
2022 QoE-oriented Adaptive Video Streaming with Edge-Client Collaborative Super-Resolution
abstract
In mobile video streaming, the ever-increasing user expectations for Quality of Experience (QoE) have prompted the integration of video super-resolution and adaptive bitrate techniques on either the mobile device or the edge server. By reconstructing high-resolution frames from low-resolution frames that have been downloaded, both high video quality and a short rebuffer time can be enjoyed. However, the exiting methods merely leverage the computing resources of the edge server or mobile device, leaving significant room for further QoE improvement. In this paper, we present an adaptive Video Streaming system with Edge-Client collaborative Super-resolution, named VSECS, to enhance users' QoE by simultaneously utilizing the computing resources of both the edge server and mobile device to reconstruct high-resolution frames collaboratively. First, we deploy a large-scale super-resolution model on the edge server and a lightweight model on the mobile device. Then, we exploit the Asynchronous Advantage Actor-Critic (A3C) algorithm to make decisions regarding the download resolution, the reconstructed target resolution, and the workload share of the mobile device, considering the network bandwidth, computing resources, and reconstruction complexity of video tiles. Furthermore, we utilize the branching actor network to enable the agent to converge to good policy stably. Trace-driven simulations on real-world bandwidth traces demonstrate that our approach can improve QoE by up to 10% compared to the state-of-the-art video streaming solutions.
Xilai Liu, Zhihui Ke, Xiaobo Zhou 0003, Tie Qiu 0001, Keqiu Li
GLOBECOM2