EDBT 2026 Demo / reviewers in the wild / expert
Jiangkai Wu
dblp:272/5246
· DBLP profile ↗
13ranked-venue papers
5as first author
12since 2021 · last 2026
0009-0007-7628-6673ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 8 · 3 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 4 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Promptus: Can Prompt Streaming Replace Video StreamingabstractWith the exponential growth of video traffic, traditional video streaming systems are approaching their limits in communication capacity. To further reduce bitrate while maintaining quality, we propose Promptus, a disruptive semantic communication system that streams prompts instead of videos. Promptus represents the real-world video with a series of "prompts" for delivery and employs Stable Diffusion to generate the same video at the receiver. To ensure that the generated video is pixel-aligned with the original video, a gradient descent-based prompt fitting framework is proposed. Further, a low-rank decomposition-based bitrate control algorithm is introduced to achieve adaptive bitrate. For inter-frame compression, an interpolation-aware fitting algorithm is proposed. Evaluations across various video genres demonstrate that, compared to H.265, Promptus can achieve more than a 4x bandwidth reduction while preserving the same perceptual quality. On the other hand, at extremely low bitrates, Promptus can enhance the perceptual quality by 0.139 and 0.118 (in LPIPS) compared to VAE and H.265, respectively, and decreases the ratio of severely distorted frames by 89.3% and 91.7%. Our work opens up a new paradigm for efficient video communication. Jiangkai Wu, Yunpeng Tan, Junlin Hao, Xinggong Zhang |
AAAI | 1 |
| 2026 | Smaller is Better: Generative Models Can Power Short Video Preloading
Jiangkai Wu, Xinggong Zhang |
ICC | 2 |
| 2026 | HybridPrompt: Bridging Generative Priors and Traditional Codecs for Mobile StreamingabstractIn Video on Demand (VoD) scenarios, traditional codecs are the industry standard due to their high decoding efficiency. However, they suffer from severe quality degradation under low bandwidth conditions. While emerging generative neural codecs offer significantly higher perceptual quality, their reliance on heavy frame-by-frame generation makes real-time playback on mobile devices impractical. We ask: is it possible to combine the blazing-fast speed of traditional standards with the superior visual fidelity of neural approaches? We present HybridPrompt, the first generative-based video system capable of achieving real-time 1080p decoding at over 150 FPS on a commercial smartphone. Specifically, we employ a hybrid architecture that encodes Keyframes using a generative model while relying on traditional codecs for the remaining frames. A major challenge is that the two paradigms have conflicting objectives: the "hallucinated" details from generative models often misalign with the rigid prediction mechanisms of traditional codecs, causing bitrate inefficiency. To address this, we demonstrate that the traditional decoding process is differentiable, enabling an end-to-end optimization loop. This allows us to use subsequent frames as additional supervision, forcing the generative model to synthesize keyframes that are not only perceptually high-fidelity but also mathematically optimal references for the traditional codec. By integrating a two-stage generation strategy, our system outperforms pure neural baselines by orders of magnitude in speed while achieving an average LPIPS gain of 8% over traditional codecs at 200kbps. Jiangkai Wu, Haoyang Wang 0015, Peiheng Wang, Zongming Guo, Xinggong Zhang |
NOSSDAV | 2 |
| 2026 | Morphe: High-Fidelity Generative Video Streaming with Vision Foundation Model
Tianyi Gong, Zijian Cao 0007, Zixing Zhang 0009, Jiangkai Wu, Xinggong Zhang, Shuguang Cui, Fangxin Wang 0001 |
NSDI | 4 |
| 2026 | Artic: AI-oriented Real-time Communication for MLLM Video Assistant
Jiangkai Wu, Junquan Zhong, Xinggong Zhang |
SIGCOMM | 1 |
| 2025 | PromptMobile: Efficient Promptus for Low Bandwidth Mobile Video StreamingabstractTraditional video compression algorithms exhibit significant quality degradation at extremely low bitrates. Promptus emerges as a new paradigm for video streaming, substantially cutting down the bandwidth essential for video streaming. However, Promptus is computationally intensive and can not run in real-time on mobile devices. This paper presents PromptMobile, an efficient acceleration framework tailored for on-device Promptus. Specifically, we propose (1) a two-stage efficient generation framework to reduce computational cost by 8.1x, (2) a fine-grained inter-frame caching to reduce redundant computations by 16.6%, (3) system-level optimizations to further enhance efficiency. The evaluations demonstrate that compared with the original Promptus, PromptMobile achieves a 13.6x increase in image generation speed. Compared with other streaming methods, PromptMobile achives an average LPIPS improvement of 0.016 (compared with H.265), reducing 60% of severely distorted frames (compared to VQGAN). Jiangkai Wu, Haoyang Wang 0015, Peiheng Wang, Xinggong Zhang, Zongming Guo |
APNet | 2 |
| 2025 | Chat with AI: The Surprising Turn of Real-time Video Communication from Human to AIabstractAI Video Chat emerges as a new paradigm for Real-time Communication (RTC), where one peer is not a human, but a Multimodal Large Language Model (MLLM). This makes interaction between humans and AI more intuitive, as if chatting face-to-face with a real person. However, this poses significant challenges to latency, because the MLLM inference takes up most of the response time, leaving very little time for video streaming. Due to network uncertainty, transmission latency becomes a critical bottleneck preventing AI from being like a real person. To address this, we call for AI-oriented RTC research, exploring the network requirement shift from "humans watching video" to "AI understanding video". We begin by recognizing the main differences between AI Video Chat and traditional RTC. Then, through prototype measurements, we identify that ultra-low bitrate is a key factor for low latency. To reduce bitrate dramatically while maintaining MLLM accuracy, we propose Context-Aware Video Streaming that recognizes the importance of each video region for chat and allocates bitrate almost exclusively to chat-important regions. To evaluate the impact of video streaming quality on MLLM accuracy, we build the first benchmark, named Degraded Video Understanding Benchmark (DeViBench). Finally, we discuss some open questions and ongoing solutions for AI Video Chat. DeViBench is open-sourced at: https://github.com/pku-netvideo/DeViBench. Jiangkai Wu, Xinggong Zhang |
HotNets | 1 |
| 2025 | 3DGCoding: Novel Framework for 3D Gaussian Video Incremental Training and CodingabstractFree-viewpoint videos (FVVs) streaming plays a pivotal role in driving the progress of immersive AR/VR applications. To be streamable, FVVs should simultaneously satisfy perframe decoding, compactness, and real-time rendering. Neural radiance based works suffer from low rendering speed, while previous 3D Gaussians (3DGs) based works necessitate large data volumes and pose challenges for compact representation. To support FVVs streaming, we present 3DGCoding, a novel framework for efficient reconstruction and coding of 3DGs for dynamic scenes. We design a learnable 3D mask to partition 3DGs into static and dynamic components, capturing all types of changes and making 3DGs quantitatively and attributively consistent for coding. Then, we propose a 2D-based coding process to further address inter-frame redundancy, achieving independent decoding and smaller data volumes. Experiments confirm that 3DGCoding could stream with 0.49MB per frame, and achieves competitive performance compared with state-of-the-art FVVs streaming methods. Peiheng Wang, Haodan Zhang, Quanlu Jia, Jiangkai Wu, Haoyang Wang 0015, Xinggong Zhang |
ICME | 4 |
| 2025 | Sync5D: Novel View Synthesis from a Single Image with 5D Consistency
Junlin Hao, Yunpeng Tan, Jiangkai Wu, Peiheng Wang, Xinggong Zhang, Zongming Guo |
PRCV (10) | 3 |
| 2024 | SRFC: Scalable Radiance Fields Streaming with Planar CodecabstractVolumetric videos afford comprehensive and immer-sive viewing experiences with six degrees of freedom (6DoF) for navigation, allowing users to move freely within a three-dimensional space. Radiance fields (RF) is emerging to reproduce photorealistic 3D scenes, which achieves lighting consistency and realistic transfer between the real and virtual worlds. In this paper, we investigate how to deliver photorealistic volumetric video under the network bandwidth constraints and low-power device. We design a novel Scalable Radiance Fields Video Streaming, SRFC, to enable streaming RF video with planar codec. An Orthogonal Layered Depth Image (OLDI) mapping is introduced to map 3D to view-dependent 2D plane images, in order to reduce bitrates and decoding complexity by the commercial 2D codec. Moreover, to address the issue of deviation of the user viewpoint, we propose a dual-tiered structure in radiance fields and optimizes user's perceptual quality by adaptive bitrate and viewpoint decision. The evaluations demonstrate that SRFC is capable of reducing bandwidth requirements by 4 × and increasing decoding frame rate by nearly 3 × while maintaining satisfactory user perception of video quality. Quanlu Jia, Haodan Zhang, Haoyang Wang 0015, Jiangkai Wu, Xinggong Zhang, Zongming Guo |
ICC | 5 |
| 2023 | ZGaming: Zero-Latency 3D Cloud Gaming by Image PredictionabstractIn cloud gaming, interactive latency is one of the most important factors in users' experience. Although the interactive latency can be reduced through typical network infrastructures like edge caching and congestion control, the interactive latency of current cloud-gaming platforms is still far from users' satisfaction. Jiangkai Wu, Yu Guan 0005, Qi Mao 0002, Yong Cui 0001, Zongming Guo, Xinggong Zhang |
SIGCOMM | 1 |
| 2023 | ActRay: Online Active Ray Sampling for Radiance FieldsabstractThanks to the high-quality reconstruction and photorealistic rendering, the Neural Radiance Field (NeRF) has garnered extensive attention and has been continuously improved. Despite its high visual quality, the prohibitive training time limits its practical application. Although significant acceleration has been achieved, it is still far from real-time training, due to the need for tens of thousands of iterations. In this paper, a feasible solution is to reduce the number of required iterations by always training the rays with the highest loss values, instead of the traditional method of training each ray with a uniform probability. To this end, we propose an online active ray sampling strategy, ActRay. Specifically, to avoid the substantial overhead of calculating the actual loss values for all rays in each iteration, a rendering-gradient-based loss propagation algorithm is presented to efficiently estimate the loss values. To further narrow the gap between the estimated loss and the actual loss, an online learning algorithm based on the Upper Confidence Bound (UCB) is proposed to control the sampling probability of the rays, thereby compensating for the bias in loss estimation. We evaluate ActRay on both real-world and synthetic scenes, and the promising results show that it accelerates radiance field training by 6.5x. Besides, we test ActRay under all kinds of radiance field representations (implicit, explicit, and hybrid), proving that it is general and effective to different representations. We believe this work will contribute to the practical application of radiance fields, because it has taken a step closer to real-time radiance field training. ActRay is open-source at: https://pku-netvideo.github.io/actray/. Jiangkai Wu, Yunpeng Tan, Quanlu Jia, Haodan Zhang, Xinggong Zhang |
SIGGRAPH Asia | 1 |
| 2020 | Identity-Aware Attribute Recognition via Real-Time Distributed Inference in Mobile Edge CloudsabstractWith the development of deep learning technologies, attribute recognition and person re-identification (re-ID) have attracted extensive attention and achieved continuous improvement via executing computing-intensive deep neural networks in cloud datacenters. However, the datacenter deployment cannot meet the real-time requirement of attribute recognition and person re-ID, due to the prohibitive delay of backhaul networks and large data transmissions from cameras to datacenters. A feasible solution thus is to employ mobile edge clouds (MEC) within the proximity of cameras and enable distributed inference. Zichuan Xu, Jiangkai Wu, Qiufen Xia, Pan Zhou 0001, Jiankang Ren, Huizhi Liang 0001 |
ACM Multimedia | 2 |