VLDB 2026 Research / reviewers in the wild / expert
Kaiyuan Hu
dblp:255/9117
· DBLP profile ↗
12ranked-venue papers
6as first author
12since 2021 · last 2026
0009-0003-2709-457XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 4 first-author · 7 since 2021Computer networks · 4 · 2 first-author · 4 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | GraVQR: A Self-Refining GNN-LLM Framework for Spatio-Temporal Traffic Prediction
Shengyi Ding, Kaiyuan Hu, Tianyu Pang, Lei Peng 0002, Ying Cui 0001, Yang Yang 0001 |
ICC | 2 |
| 2026 | Privis: Towards Content-Aware Secure Volumetric Video Delivery
Kaiyuan Hu, Hong Kang, Yili Jin 0001, Junhua Liu 0003, Chengming Hu, Haolun Wu, Xue (Steve) Liu |
ICC | 1 |
| 2026 | Lightweight Edge-Deployable Voltage Prediction Method for Distribution Networks With High Photovoltaic PenetrationabstractUnder the continued growth of high penetration distributed photovoltaics (PV), PV interface voltages exhibit pronounced randomness, rapid fluctuations, and even overvoltage, making short-term voltage prediction increasingly critical yet more challenging under dynamic operating conditions. This growing reliance on accurate and timely prediction directly exposes the shortcomings of existing cloud-based forecasting schemes, which are constrained by communication delays, bandwidth limitations, and unstable connectivity, ultimately restricting their ability to provide the fast response required by a prediction-control loop. To address these issues, this study proposes a lightweight edge-deployable voltage prediction (LEDVP) method that consist of an ensemble algorithm and a distributed framework. The ensemble algorithm combines variational mode decomposition for voltage-signal decomposition and extreme gradient boosting for extracting the most informative features, which are then fed into a Mamba state-space sequence model to perform efficient linear-time temporal inference, forming a compact and lightweight prediction process suitable for embedded edge platforms. Moreover, by quantizing model parameters, pruning redundant structures, reusing intermediate caches, and optimizing the model specifically, the LEDVP can be deployed in the distributed framework. Experimental results show that the LEDVP method achieves near-lossless accuracy with a 1–3 s latency on NVIDIA Orin and yields approximately 40–50% lower root mean square error compared with representative decomposition-based neural network methods, while the LEDVP framework maintains stable resource usage. These results demonstrate the engineering feasibility and practical value of LEDVP for real-time voltage prediction, PV interface voltage regulation, and edge-side autonomy in PV-rich distribution networks. Zhijin Lyu, Wei Liu 0063, Xibin Shan, Kaiyuan Hu |
IEEE Trans. Ind. Informatics | 6 |
| 2025 | Generative AI for Immersive Video: Recent Advances and Future OpportunitiesabstractImmersive video serves as a key component of eXtended Reality (XR) that aims to create and interact with simulated virtual or hybrid environments. Such a technology allows users to experience immersive sensations that transcend time and space, and meanwhile continuously providing training data for emerging technologies like Embodied AI. Thanks to the advancements in sensing, computing, and display, recent years have witnessed many excellent works for XR and related hardware or software systems. However, challenges like high creation cost, lack of immersion, and limited scalability hinder the practical application of immersive video services. Whilst recently emerged generative artificial intelligence (GenAI) provides us with new insights in tackling existing challenges. In this paper, we conduct a comprehensive survey into the recent advances and future opportunities on how GenAI can benefit immersive video services. By introducing a systematic taxonomy, we meticulously classify the pertinent techniques and applications into three well-defined categories aligned with the pipeline of immersive video service: content creation, network delivery, and client-side display. This categorization enables a structured exploration of the diverse roles on how GenAI can benefit immersive video service, providing a framework for a more comprehensive understanding and evaluation of these technologies. To the best of our knowledge, this work is the first systematic survey of GenAI in XR settings, laying a foundation for future research in this interdisciplinary domain. Kaiyuan Hu, Yili Jin 0001, Hao Zhou 0013, Linfeng Du, Jiangchuan Liu |
IJCAI | 1 |
| 2025 | LiveVV: Human-Centered Live Volumetric Video Streaming SystemabstractVolumetric video (VV) has emerged as a prominent medium within the realm of extended reality (XR) with advancements in computer graphics and depth capture hardware. Users can fully immersive themselves in VV with the ability to switch their viewport in six degree of freedom (DOF), including three rotational dimensions (yaw, pitch, and roll) and three translational dimensions (X, Y, and Z). Different from traditional 2-D videos that are composed of pixel matrices, VVs employ point clouds, meshes, or voxels to represent a volumetric scene, resulting in significantly larger data sizes. While previous works have successfully achieved VV streaming in video-on-demand scenarios, the live streaming of VV remains an unresolved challenge due to the limited network bandwidth and stringent latency constraints. In this article, we proposeLiveVV, a holistic live VV streaming system that integrates multiview capture, scene segmentation and reuse, adaptive transmission, and real-time rendering.LiveVVfeatures lightweight VV capture modules for easy deployment, processes static and dynamic content separately to reduce bandwidth consumption, and incorporates a VV adaptive bitrate streaming algorithm (VABR) to ensure fluent playback with high-quality experience. Real-world implementation and evaluation demonstrate thatLiveVVachieves live VV streaming at 24 FPS frame rate with less than 350-ms latency on average, meeting the requirements of real-life application. Kaiyuan Hu, Yongting Chen, Kaiying Han, Yili Jin 0001, Junhua Liu 0003, Fangxin Wang 0001 |
IEEE Internet Things J. | 1 |
| 2025 | 3D Video Conferencing via On-Hand DevicesabstractVideo conferencing has become indispensable in human communication. Researchers are exploring immersive capabilities to enhance video conferencing experiences by delivering realistic interactions. However, existing methods have stringent and extra hardware beyond a typical video conference, including multiple depth cameras, large screens, and headsets, which pose obstacles to the widespread adoption due to high costs and complex setups. Thus, there is an urgent demand for light-weight systems using only on-hand devices including single RGB camera and standard screen, without additional hardware. We propose DVCO, a novel 3D video conferencing system via on-hand devices. With DVCO, users can experience lifelike virtual conferencing that includes natural contact and interactive features. To achieve this, DVCO has two main components. Virtual Camera Transformation (VCT) and New View Generator (NVG). VCT computes a downscaled sender image from tracking to determine viewpoint and gaze vector, enhancing virtual presence on standard screens. NVG takes an input frame and desired view angle to produce an output reflecting the new view from a single RGB camera. Together, these provide an affordable, easy-to-integrate enhancement for current video conferencing systems without expensive upgrades. Through a user study, it has been demonstrated that DVCO offers an exceptional level of immersion when compared to traditional systems. Experiments are conducted to showcase the superior performance of VCT and NVG in comparison to baseline methods. Yili Jin 0001, Xize Duan, Kaiyuan Hu, Fangxin Wang 0001, Xue (Steve) Liu |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | Video Conferencing With Predictive Generation and Collaborative Computation Across Mobile HeadsetsabstractVirtual Reality (VR) has emerged as a transformative platform for remote collaboration, but its adoption for video conferencing is hindered by challenges related to facial expression reconstruction and computational resource constraints, especially on economical mobile VR headsets. This paper introduces a novel system for VR video conferencing that addresses these challenges through two key modules: Predictive Generation and Collaborative Computation. Predictive Generation leverages multimodal inputs, including voice, head motion, and eye blinks, to synthesize realistic facial animations with low latency, eliminating the need for high-precision hardware. Collaborative Computation enhances computational efficiency by employing a game-theoretic framework for resource sharing among users. Experimental evaluations demonstrate that our system delivers immersive and realistic VR video conferencing experiences with superior facial expression reconstruction and efficient resource utilization. Our approach makes VR video conferencing more accessible and practical for a broader audience across mobile headsets. Yili Jin 0001, Xize Duan, Kaiyuan Hu, Fangxin Wang 0001, Xue (Steve) Liu, Jiangchuan Liu |
IEEE Trans. Mob. Comput. | 3 |
| 2024 | MMCOUNT: Stationary Crowd Counting System Based on Commodity Millimeter-Wave RadarabstractMillimeter wave sensing promises the capability of sensing the surrounding moving people. However, it is still challenging for stationary crowds because objects with few motions (like changing sitting position) are easily treated as a cluster of noise and thus neglected. In this paper, we propose that people’s respiration and natural fidgeting (restless behavior) carry valuable information, which could be captured by millimeter (mmWave) radar. By performing processing on the captured data including signal enhancement and object recognition, we can successfully extract the number of a crowd and the position of each individual. To verify our system, we test it in different locations like hall, classroom, and meeting room to simulate different practical scenarios including watching a movie, having a class, or attending a meeting. The evaluation results show that our proposed approach could reach a high counting accuracy of up to 95.8% even at small separation distances of 0.4m. Kaiyuan Hu, Hongjie Liao, Fangxin Wang 0001 |
ICASSP | 1 |
| 2024 | AI-Enhanced Virtual Reality in Medicine: A Comprehensive Survey
Kaiyuan Hu, Danny Ziyi Chen, Jian Wu 0001 |
IJCAI | 2 |
| 2024 | TeleOR: Real-Time Telemedicine System for Full-Scene Operating Room
Kaiyuan Hu, Qian Shao, Jintai Chen, Danny Ziyi Chen, Jian Wu 0001 |
MICCAI (6) | 2 |
| 2023 | Understanding User Behavior in Volumetric Video Watching: Dataset, Analysis and PredictionabstractVolumetric video emerges as a new attractive video paradigm in recent years since it provides an immersive and interactive 3D viewing experience with six degree-of-freedom (DoF). Unlike traditional 2D or panoramic videos, volumetric videos require dense point clouds, voxels, meshes, or huge neural models to depict volumetric scenes, which results in a prohibitively high bandwidth burden for video delivery. Users' behavior analysis, especially the viewport and gaze analysis, then plays a significant role in prioritizing the content streaming within users' viewport and degrading the remaining content to maximize user QoE with limited bandwidth. Although understanding user behavior is crucial, to the best of our best knowledge, there are no available 3D volumetric video viewing datasets containing fine-grained user interactivity features, not to mention further analysis and behavior prediction. Kaiyuan Hu, Yili Jin 0001, Junhua Liu 0003, Yongting Chen, Miao Zhang 0003, Fangxin Wang 0001 |
ACM Multimedia | 1 |
| 2023 | FSVVD: A Dataset of Full Scene Volumetric VideoabstractRecent years have witnessed a rapid development of immersive multimedia which bridges the gap between the real world and virtual space. Volumetric videos, as an emerging representative 3D video paradigm that empowers extended reality, stand out to provide unprecedented immersive and interactive video watching experience. Despite the tremendous potential, the research towards 3D volumetric video is still in its infancy, relying on sufficient and complete datasets for further exploration. However, existing related volumetric video datasets mostly only include a single object, lacking details about the scene and the interaction between them. In this paper, we focus on the current most widely used data format, point cloud, and for the first time release a full-scene volumetric video dataset that includes multiple people and their daily activities interacting with the external environments. Comprehensive dataset description and analysis are conducted, with potential usage of this dataset. The dataset and additional tools can be accessed via the following website: https://cuhksz-inml.github.io/full_scene_volumetric_video_dataset/. Kaiyuan Hu, Yili Jin 0001, Junhua Liu 0003, Fangxin Wang 0001 |
MMSys | 1 |