VLDB 2026 Research / reviewers in the wild / expert
Yili Jin 0001
dblp:60/10595-1
· DBLP profile ↗
22ranked-venue papers
9as first author
22since 2021 · last 2026
0000-0002-7127-8902ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 13 · 6 first-author · 13 since 2021Computer networks · 10 · 4 first-author · 10 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Privis: Towards Content-Aware Secure Volumetric Video Delivery
Kaiyuan Hu, Hong Kang, Yili Jin 0001, Junhua Liu 0003, Chengming Hu, Haolun Wu, Xue (Steve) Liu |
ICC | 3 |
| 2026 | Self-Supervised Compression and Artifact Correction for Streaming Underwater Imaging SonarabstractReal-time imaging sonar is crucial for underwater monitoring where optical sensing fails, but its use is limited by low uplink bandwidth and severe sonar-specific artifacts (speckle, motion blur, reverberation, acoustic shadows) affecting up to 98% of frames. We present SCOPE, a self-supervised framework that jointly performs compression and artifact correction without clean–noise pairs or synthetic assumptions. SCOPE combines (i) Adaptive Codebook Compression (ACC), which learns frequency-encoded latent representations tailored to imaging sonar, with (ii) Frequency-Aware Multiscale Segmentation (FAMS), which decomposes frames into low-frequency structure and sparse high-frequency dynamics while suppressing rapidly fluctuating artifacts. A hedging training strategy further guides frequency-aware learning using low-pass proxy pairs generated without labels. Evaluated on months of in-situ ARIS sonar data, SCOPE achieves a structural similarity index (SSIM) of 0.77, representing a 40% improvement over prior self-supervised denoising baselines, at bitrates down to ≤ 0.0118 bpp. It reduces uplink bandwidth by more than 80% while improving downstream detection. The system runs in real time, with 3.1 ms encoding on an embedded GPU and 97 ms full multi-layer decoding on the server end. SCOPE has been deployed for months in three Pacific Northwest rivers to support real-time salmon enumeration and environmental monitoring in the wild. Results demonstrate that learning frequency-structured latents enables practical, low-bitrate sonar streaming with preserved signal details under real-world deployment conditions. Rongsheng Qian, Chi Xu 0004, Xiaoqiang Ma, Hao Fang 0012, Yili Jin 0001, William I. Atlas, Jiangchuan Liu |
WACV | 5 |
| 2025 | Generative AI for Immersive Video: Recent Advances and Future OpportunitiesabstractImmersive video serves as a key component of eXtended Reality (XR) that aims to create and interact with simulated virtual or hybrid environments. Such a technology allows users to experience immersive sensations that transcend time and space, and meanwhile continuously providing training data for emerging technologies like Embodied AI. Thanks to the advancements in sensing, computing, and display, recent years have witnessed many excellent works for XR and related hardware or software systems. However, challenges like high creation cost, lack of immersion, and limited scalability hinder the practical application of immersive video services. Whilst recently emerged generative artificial intelligence (GenAI) provides us with new insights in tackling existing challenges. In this paper, we conduct a comprehensive survey into the recent advances and future opportunities on how GenAI can benefit immersive video services. By introducing a systematic taxonomy, we meticulously classify the pertinent techniques and applications into three well-defined categories aligned with the pipeline of immersive video service: content creation, network delivery, and client-side display. This categorization enables a structured exploration of the diverse roles on how GenAI can benefit immersive video service, providing a framework for a more comprehensive understanding and evaluation of these technologies. To the best of our knowledge, this work is the first systematic survey of GenAI in XR settings, laying a foundation for future research in this interdisciplinary domain. Kaiyuan Hu, Yili Jin 0001, Hao Zhou 0013, Linfeng Du, Jiangchuan Liu |
IJCAI | 2 |
| 2025 | Exploring Multimodal Foundation AI and Expert-in-the-Loop for Sustainable Management of Wild Salmon Fisheries in Indigenous RiversabstractWild salmon are essential to the ecological, economic, and cultural sustainability of the North Pacific Rim. Yet climate variability, habitat loss, and data limitations in remote ecosystems that lack basic infrastructure support pose significant challenges to effective fisheries management. This project explores the integration of multimodal foundation AI and expert-in-the-loop frameworks to enhance wild salmon monitoring and sustainable fisheries management in Indigenous rivers across Pacific Northwest. By leveraging video and sonar-based monitoring, we develop AI-powered tools for automated species identification, counting, and length measurement, reducing manual effort, expediting delivery of results, and improving decision-making accuracy. Expert validation and active learning frameworks ensure ecological relevance while reducing annotation burdens. To address unique technical and societal challenges, we bring together a cross-domain, interdisciplinary team of university researchers, fisheries biologists, Indigenous stewardship practitioners, government agencies, and conservation organizations. Through these collaborations, our research fosters ethical AI co-development, open data sharing, and culturally informed fisheries management. Chi Xu 0004, Yili Jin 0001, Sami Ma, Rongsheng Qian, Hao Fang 0012, Jiangchuan Liu, Xue (Steve) Liu, Edith C. H. Ngai, William I. Atlas, Katrina M. Connors, Mark A. Spoljaric |
IJCAI | 2 |
| 2025 | Generative AI for Multimedia Communication: Recent Advances, An Information-Theoretic Framework, and Future OpportunitiesabstractRecent breakthroughs in generative artificial intelligence (AI) are transforming multimedia communication. This paper systematically reviews key recent advancements across generative AI for multimedia communication, emphasizing transformative models like diffusion and transformers. However, conventional information-theoretic frameworks fail to address semantic fidelity, critical to human perception. We propose an innovative semantic information-theoretic framework, introducing semantic entropy, mutual information, channel capacity, and rate-distortion concepts specifically adapted to multimedia applications. This framework redefines multimedia communication from purely syntactic data transmission to semantic information conveyance. We further highlight future opportunities and critical research directions. We chart a path toward robust, efficient, and semantically meaningful multimedia communication systems by bridging generative AI innovations with information theory. This exploratory paper aims to inspire a semantic-first paradigm shift, offering a fresh perspective with significant implications for future multimedia research. Yili Jin 0001, Xue (Steve) Liu, Jiangchuan Liu |
ACM Multimedia | 1 |
| 2025 | Generative Flow Networks for Personalized Multimedia Systems: A Case Study on Short Video FeedsabstractMultimedia systems underpin modern digital interactions, facilitating seamless integration and optimization of resources across diverse multimedia applications. To meet growing personalization demands, multimedia systems must efficiently manage competing resource needs, adaptive content, and user-specific data handling. This paper introduces Generative Flow Networks (GFlowNets, GFNs) as a brave new framework for enabling personalized multimedia systems. By integrating multi-candidate generative modeling with flow-based principles, GFlowNets offer a scalable and flexible solution for enhancing user-specific multimedia experiences. To illustrate the effectiveness of GFlowNets, we focus on short video feeds, a multimedia application characterized by high personalization demands and significant resource constraints, as a case study. Our proposed GFlowNet-based personalized feeds algorithm demonstrates superior performance compared to traditional rule-based and reinforcement learning methods across critical metrics, including video quality, resource utilization efficiency, and delivery cost. Moreover, we propose a unified GFlowNet-based framework generalizable to other multimedia systems, highlighting its adaptability and wide-ranging applicability. These findings underscore the potential of GFlowNets to advance personalized multimedia systems by addressing complex optimization challenges and supporting sophisticated multimedia application scenarios. Yili Jin 0001, Ling Pan, Rui-Xiao Zhang, Jiangchuan Liu, Xue (Steve) Liu |
ACM Multimedia | 1 |
| 2025 | SemConf: A System for Multiparty Semantic Video ConferencingabstractMulti-party real-time video conferencing has become an indispensable service in industrial production and daily life. However, the current dynamic and limited network resources can no longer meet the growing service demands of users, resulting lagging and low visual quality. The emerging semantic transmission, together with the network-wide redundant computation capacity, provides new opportunities towards a new paradigm of semantic video conferencing. The key challenge of such fusion lies in the interplay of traditional streaming adaptation and the new semantic processing, calling for a holistic mechanism to optimize the service provision with compatibility and efficiency. In this paper, we for the first time address this challenge, and propose SemConf, a novel framework that integrate the semantic transmission into the video conferencing towards optimal user QoE. Our extensive evaluations, against state-of-the-art baselines, reveal that SemConf achieves a substantial improvement in QoE, with up to 33.6% enhancement in bandwidth-constrained environment. Overall, this work highlights the critical role of the coordination algorithm in balancing computational load and network throughput, showcasing SemConf as a transformative approach in the realm of semantic video conferencing. Xize Duan, Yili Jin 0001, Lei Zhang 0066, Fangxin Wang 0001 |
NOSSDAV | 2 |
| 2025 | TrackerSplat: Exploiting Point Tracking for Fast and Robust Dynamic 3D Gaussians ReconstructionabstractRecent advancements in 3D Gaussian Splatting (3DGS) have demonstrated its potential for efficient and photorealistic 3D reconstructions, which is crucial for diverse applications such as robotics and immersive media. However, current Gaussian-based methods for dynamic scene reconstruction struggle with large inter-frame displacements, leading to artifacts and temporal inconsistencies under fast object motions. To address this, we introduce TrackerSplat, a novel method that integrates advanced point tracking methods to enhance the robustness and scalability of 3DGS for dynamic scene reconstruction. TrackerSplat utilizes off-the-shelf point tracking models to extract pixel trajectories and triangulate per-view pixel trajectories onto 3D Gaussians to guide the relocation, rotation, and scaling of Gaussians before training. This strategy effectively handles large displacements between frames, dramatically reducing the fading and recoloring artifacts prevalent in prior methods. By accurately positioning Gaussians prior to gradient-based optimization, TrackerSplat overcomes the quality degradation associated with large frame gaps when processing multiple adjacent frames in parallel across multiple devices, thereby boosting reconstruction throughput while preserving rendering quality. Experiments on real-world datasets confirm the robustness of TrackerSplat in challenging scenarios with significant displacements, achieving superior throughput under parallel settings and maintaining visual quality compared to baselines. The code is available at https://github.com/yindaheng98/TrackerSplat. Daheng Yin, Isaac Ding, Yili Jin 0001, Jianxin Shi 0005, Jiangchuan Liu |
SIGGRAPH Asia | 3 |
| 2025 | LiveVV: Human-Centered Live Volumetric Video Streaming SystemabstractVolumetric video (VV) has emerged as a prominent medium within the realm of extended reality (XR) with advancements in computer graphics and depth capture hardware. Users can fully immersive themselves in VV with the ability to switch their viewport in six degree of freedom (DOF), including three rotational dimensions (yaw, pitch, and roll) and three translational dimensions (X, Y, and Z). Different from traditional 2-D videos that are composed of pixel matrices, VVs employ point clouds, meshes, or voxels to represent a volumetric scene, resulting in significantly larger data sizes. While previous works have successfully achieved VV streaming in video-on-demand scenarios, the live streaming of VV remains an unresolved challenge due to the limited network bandwidth and stringent latency constraints. In this article, we proposeLiveVV, a holistic live VV streaming system that integrates multiview capture, scene segmentation and reuse, adaptive transmission, and real-time rendering.LiveVVfeatures lightweight VV capture modules for easy deployment, processes static and dynamic content separately to reduce bandwidth consumption, and incorporates a VV adaptive bitrate streaming algorithm (VABR) to ensure fluent playback with high-quality experience. Real-world implementation and evaluation demonstrate thatLiveVVachieves live VV streaming at 24 FPS frame rate with less than 350-ms latency on average, meeting the requirements of real-life application. Kaiyuan Hu, Yongting Chen, Kaiying Han, Yili Jin 0001, Junhua Liu 0003, Fangxin Wang 0001 |
IEEE Internet Things J. | 6 |
| 2025 | 3D Video Conferencing via On-Hand DevicesabstractVideo conferencing has become indispensable in human communication. Researchers are exploring immersive capabilities to enhance video conferencing experiences by delivering realistic interactions. However, existing methods have stringent and extra hardware beyond a typical video conference, including multiple depth cameras, large screens, and headsets, which pose obstacles to the widespread adoption due to high costs and complex setups. Thus, there is an urgent demand for light-weight systems using only on-hand devices including single RGB camera and standard screen, without additional hardware. We propose DVCO, a novel 3D video conferencing system via on-hand devices. With DVCO, users can experience lifelike virtual conferencing that includes natural contact and interactive features. To achieve this, DVCO has two main components. Virtual Camera Transformation (VCT) and New View Generator (NVG). VCT computes a downscaled sender image from tracking to determine viewpoint and gaze vector, enhancing virtual presence on standard screens. NVG takes an input frame and desired view angle to produce an output reflecting the new view from a single RGB camera. Together, these provide an affordable, easy-to-integrate enhancement for current video conferencing systems without expensive upgrades. Through a user study, it has been demonstrated that DVCO offers an exceptional level of immersion when compared to traditional systems. Experiments are conducted to showcase the superior performance of VCT and NVG in comparison to baseline methods. Yili Jin 0001, Xize Duan, Kaiyuan Hu, Fangxin Wang 0001, Xue (Steve) Liu |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2025 | Multi-Task Reinforcement Learning-Based Multiple Access for Dynamic Wireless NetworksabstractWith the rapid development of emergent applications, wireless networks require the provision of high throughput. Meanwhile, wireless scenarios exhibit highly dynamic characteristics, involving frequent changes in the network scale and traffic. To satisfy the high demand for new applications in dynamic wireless scenarios, a novel medium access control (MAC) protocol is required to allow stations to access the channel with high efficiency and adaptability. Based on multi-agent reinforcement learning (MARL), we propose a new MAC protocol, Multi-task Transformer-based Multiple Access (MTMA). Multi-task learning is applied to train a single actor to adapt to multiple wireless environments simultaneously. To improve the scalability, we propose a transformer-based critic network, which can scale to different wireless scenarios. Moreover, a novel network called “Generalization for N (Gen-N)” network is proposed to enhance the generalization ability. We conduct simulation experiments to demonstrate that MTMA: 1) achieves over 95% of upper bound of throughput while maximizing the fairness performance; 2) outperforms classic MAC protocol and MARL-based baselines in scenarios with saturated and light traffic; 3) can adapt to environmental changes quickly in dynamic scenarios; 4) can generalize to unseen scenarios during training. Finally, the ablation experiments are conducted to evaluate the effectiveness of components used in MTMA. Xinghua Sun, Yili Jin 0001, Fangxin Wang 0001 |
IEEE Trans. Mob. Comput. | 3 |
| 2025 | Video Conferencing With Predictive Generation and Collaborative Computation Across Mobile HeadsetsabstractVirtual Reality (VR) has emerged as a transformative platform for remote collaboration, but its adoption for video conferencing is hindered by challenges related to facial expression reconstruction and computational resource constraints, especially on economical mobile VR headsets. This paper introduces a novel system for VR video conferencing that addresses these challenges through two key modules: Predictive Generation and Collaborative Computation. Predictive Generation leverages multimodal inputs, including voice, head motion, and eye blinks, to synthesize realistic facial animations with low latency, eliminating the need for high-precision hardware. Collaborative Computation enhances computational efficiency by employing a game-theoretic framework for resource sharing among users. Experimental evaluations demonstrate that our system delivers immersive and realistic VR video conferencing experiences with superior facial expression reconstruction and efficient resource utilization. Our approach makes VR video conferencing more accessible and practical for a broader audience across mobile headsets. Yili Jin 0001, Xize Duan, Kaiyuan Hu, Fangxin Wang 0001, Xue (Steve) Liu, Jiangchuan Liu |
IEEE Trans. Mob. Comput. | 1 |
| 2025 | Neural Image Compression with Regional DecodingabstractAs advancements are made in technology such as AR/VR and high-resolution photography, there is a growing need for a function in image compression named regional decoding . This function lets an image be encoded as a whole, but allows for an arbitrary region to be decoded using only a small part of the bitstream. However, existing neural image compression methods lack support for this crucial functionality. In this article, we propose a novel approach called the slicing en/decoder , which addresses the need for regional decoding while maintaining performance on par with state-of-the-art methods. Our approach is based on the insight that, during the compression process, local information within pixels holds greater importance than global information. By leveraging this understanding, we divide the image into different bitstreams according to cross-boundary patterns. Consequently, for a selected region, our method can intelligently choose specific portions of the bitstreams to decode only that particular region of interest. Furthermore, we extend the application of our method to 360° image compression, allowing for efficient encoding and decoding of immersive visual content. Moreover, our proposed technique offers the capability to decode regions identically, which paves the way for future advancements in regional video decoding. Our experimental results demonstrate that our method maintains performance on par with state-of-the-art methods while providing the functionality of regional decoding . In conclusion, this article presents a significant step forward in image compression technology, offering enhanced flexibility and efficiency for emerging applications in digital media. Yili Jin 0001, Jiahao Li 0001, Bin Li 0012, Yan Lu 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2024 | HeadsetOff: Enabling Photorealistic Video Conferencing on Economical VR HeadsetsabstractVirtual Reality (VR) has become increasingly popular for remote collaboration, but video conferencing poses challenges when the user's face is covered by the headset. Existing solutions have limitations in terms of accessibility. In this paper, we propose HeadsetOff, a novel system that achieves photorealistic video conferencing on economical VR headsets by leveraging voice-driven face reconstruction. HeadsetOff consists of three main components: a multimodal predictor, a generator, and an adaptive controller. The predictor effectively predicts user future behavior based on different modalities. The generator employs voice, head motion, and eye blink to animate the human face. The adaptive controller dynamically selects the appropriate generator model based on the trade-off between video quality and delay. Experimental results demonstrate the effectiveness of HeadsetOff in achieving high-quality, low-latency video conferencing on economical VR headsets. Yili Jin 0001, Xize Duan, Fangxin Wang 0001, Xue (Steve) Liu |
ACM Multimedia | 1 |
| 2024 | Privacy-Preserving Gaze-Assisted Immersive Video StreamingabstractImmersive videos, also known as 360$^{\circ }$videos, have gained significant attention in recent years due to their ability to provide an interactive and engaging experience. However, the development of immersive video streaming faces several challenges, including privacy concerns, the need for accurate viewport prediction, and efficient bandwidth allocation. In this paper, we propose a comprehensive system that integrates three specialized modules: the Privacy Protection module, the Viewport Prediction module, and the Bitrate Allocation module. The Privacy Protection module introduces a novel approach to differential privacy tailored for immersive video environments, considering the spatial and temporal correlations in viewport and gaze motion data. The Viewport Prediction module leverages a crossmodal attention mechanism based on the transformer to predict user viewport movements by analyzing the complex interactions between historical data, video content, and gaze patterns. The Bitrate Allocation module employs an adaptive tile-based bitrate allocation strategy using an exponential decay function to optimize video quality and maximize user quality of experience. Experimental results demonstrate that our proposed framework outperforms three state-of-the-art integrated frameworks, achieving an average QoE improvement of 21.61%. This paper offers substantial novelty in addressing privacy concerns, leveraging gaze information for viewport prediction, and utilizing underlying correlations between different features. Yili Jin 0001, Wenyi Morty Zhang, Fangxin Wang 0001, Xue (Steve) Liu |
IEEE Trans. Mob. Comput. | 1 |
| 2023 | When Configuration Verification Meets Machine Learning: A DRL Approach for Finding Minimum k-Link Failures
Yili Jin 0001, Lizhao You, Liqun Fu 0001, Qiao Xiang |
APNOMS | 2 |
| 2023 | Understanding User Behavior in Volumetric Video Watching: Dataset, Analysis and PredictionabstractVolumetric video emerges as a new attractive video paradigm in recent years since it provides an immersive and interactive 3D viewing experience with six degree-of-freedom (DoF). Unlike traditional 2D or panoramic videos, volumetric videos require dense point clouds, voxels, meshes, or huge neural models to depict volumetric scenes, which results in a prohibitively high bandwidth burden for video delivery. Users' behavior analysis, especially the viewport and gaze analysis, then plays a significant role in prioritizing the content streaming within users' viewport and degrading the remaining content to maximize user QoE with limited bandwidth. Although understanding user behavior is crucial, to the best of our best knowledge, there are no available 3D volumetric video viewing datasets containing fine-grained user interactivity features, not to mention further analysis and behavior prediction. Kaiyuan Hu, Yili Jin 0001, Junhua Liu 0003, Yongting Chen, Miao Zhang 0003, Fangxin Wang 0001 |
ACM Multimedia | 3 |
| 2023 | FSVVD: A Dataset of Full Scene Volumetric VideoabstractRecent years have witnessed a rapid development of immersive multimedia which bridges the gap between the real world and virtual space. Volumetric videos, as an emerging representative 3D video paradigm that empowers extended reality, stand out to provide unprecedented immersive and interactive video watching experience. Despite the tremendous potential, the research towards 3D volumetric video is still in its infancy, relying on sufficient and complete datasets for further exploration. However, existing related volumetric video datasets mostly only include a single object, lacking details about the scene and the interaction between them. In this paper, we focus on the current most widely used data format, point cloud, and for the first time release a full-scene volumetric video dataset that includes multiple people and their daily activities interacting with the external environments. Comprehensive dataset description and analysis are conducted, with potential usage of this dataset. The dataset and additional tools can be accessed via the following website: https://cuhksz-inml.github.io/full_scene_volumetric_video_dataset/. Kaiyuan Hu, Yili Jin 0001, Junhua Liu 0003, Fangxin Wang 0001 |
MMSys | 2 |
| 2023 | CaV3: Cache-assisted Viewport Adaptive Volumetric Video StreamingabstractVolumetric video (VV) recently emerges as a new form of video application providing a photorealistic immersive 3D viewing experience with 6 degree-of-freedom (DoF), which empowers many applications such as VR, AR, and Metaverse. A key problem therein is how to stream the enormous size VV through the network with limited bandwidth. Existing works mostly focused on predicting the viewport for a tiling-based adaptive VV streaming, which however only has quite a limited effect on resource saving. We argue that the content repeatability in the viewport can be further leveraged, and for the first time, propose a client-side cache-assisted strategy that aims to buffer the repeatedly appearing VV tiles in the near future so as to reduce the redundant VV content transmission. The key challenges exist in three aspects, including (1) feature extraction and mining in 6 DoF VV context, (2) accurate long-term viewing pattern estimation and (3) optimal caching scheduling with limited capacity. In this paper, we propose CaV3, an integrated cache-assisted viewport adaptive VV streaming framework to address the challenges. CaV3 employs a Long-short term Sequential prediction model (LSTSP) that achieves accurate short-term, mid-term and long-term viewing pattern prediction with a multi-modal fusion model by capturing the viewer's behavior inertia, current attention, and subjective intention. Besides, CaV3 also contains a contextual MAB-based caching adaptation algorithm (CCA) to fully utilize the viewing pattern and solve the optimal caching problem with a proved upper bound regret. Compared to existing VV datasets only containing single or co-located objects, we for the first time collect a comprehensive dataset with sufficient practical unbounded 360° scenes. The extensive evaluation of the dataset confirms the superiority of CaV3, which outperforms the SOTA algorithm by 15.6%-43% in viewport prediction and 13%-40% in system utility. Junhua Liu 0003, Boxiang Zhu, Fangxin Wang 0001, Yili Jin 0001, Shuguang Cui |
VR | 4 |
| 2023 | Ebublio: Edge-Assisted Multiuser 360° Video StreamingabstractAs one of the most important manifestations of virtual reality (VR), 360° panoramic videos in recent years have experienced booming development due to the desire for immersive and interactive experiences. Compared to traditional videos, 360° videos are featured with uncertain user Field of View (FoV), more sensitive delay tolerance, and much higher bandwidth requirement, bringing unprecedented challenges to 360° video streaming. Meanwhile, the development of 5G and mobile edge computing starts to pave the way for high-bandwidth low-latency video streaming. Some preliminary works focus on either individual FoV prediction or multiuser Quality of Experience (QoE) oriented cache strategy design, while how to design a holistic solution toward optimizing the overall user QoE with considerations over fairness and long-term system cost remains a nontrivial problem. In this article, we proposeEbublio, a novel intelligent edge caching framework to address the aforementioned challenges in 360° video streaming.Ebublioconsists of a collaborative FoV prediction (CFP) module and a long-term tile caching optimization (LTO) module to jointly optimize the long-term user QoE and system cost. The former module integrates the features of video content, user trajectory, and other users’ records for combined prediction. The latter one employs the Lyapunov framework and a subgradient optimization approach toward the optimal caching replacement policy. Our trace-driven evaluation demonstrates the superiority of our framework, with about 42% improvement in FoV prediction, and 36% improvement in QoE at similar traffic consumption. Yili Jin 0001, Junhua Liu 0003, Fangxin Wang 0001, Shuguang Cui |
IEEE Internet Things J. | 1 |
| 2022 | Where Are You Looking?: A Large-Scale Dataset of Head and Gaze Behavior for 360-Degree Videos and a Pilot Studyabstract360° videos in recent years have experienced booming development. Compared to traditional videos, 360° videos are featured with uncertain user behaviors, bringing opportunities as well as challenges. Datasets are necessary for researchers and developers to explore new ideas and conduct reproducible analyses for fair comparisons among different solutions. However, existing related datasets mostly focused on users' field of view (FoV), ignoring the more important eye gaze information, not to mention the integrated extraction and analysis of both FoV and eye gaze. Besides, users' behavior patterns are highly related to videos, yet most existing datasets only contained videos with subjective and qualitative classification from video genres, which lack quantitative analysis and fail to characterize the intrinsic properties of a video scene. To this end, we first propose a quantitative taxonomy for 360° videos that contains three objective technical metrics. Based on this taxonomy, we collect a dataset containing users' head and gaze behaviors simultaneously, which outperforms existing datasets with rich dimensions, large scale, strong diversity, and high frequency. Then we conduct a pilot study on users' behaviors and get some interesting findings such as user's head direction will follow his/her gaze direction with the most possible time interval. A case of application in tile-based 360° video streaming based on our dataset is later conducted, demonstrating a great performance improvement of existing works by leveraging our provided gaze information. Our dataset is available at https://cuhksz-inml.github.io/head_gaze_dataset/ Yili Jin 0001, Junhua Liu 0003, Fangxin Wang 0001, Shuguang Cui |
ACM Multimedia | 1 |
| 2022 | Fast Configuration Change Impact Analysis for Network Overlay Data Center NetworksabstractThis paper presents the first network configuration verifier that provides fast all-pair reachability analysis of incremental configuration changes for network overlay data center networks (DCNs). Network overlay DCNs leverage distributed routing protocol on edge leaf switches to disseminate overlay routes and establish overlay tunnels. In addition, network overlay DCNs use access control lists, microsegmentation policy, policy-based routing and firewall policy to control east-west and north-south traffic. Although some incremental verification approaches have been proposed, they either do not support certain forwarding features of the network, or are not efficient. Our configuration verifier addresses these issues through the following components: 1) a port predicate based forwarding model that is general to support all features; 2) fine-grained association technique to index possibly affected reachable pairs by changed interfaces in the original network; and 3) required waypoint path computation that finds all reachable pairs related to changed interfaces in the new network. Based on these components, our verifier presents two incremental verification algorithms that are specially designed for different service update cases. Experiment results show that our incremental verification algorithms are accurate and fast. For all-pair reachability, our verifier performs change-impact analysis within 15s for networks with 200 leafs (4000 subnets and 16 million pairs), outperforming existing approaches by up to 10x. Lizhao You, Jiahua Zhang 0005, Yili Jin 0001 |
IEEE/ACM Trans. Netw. | 3 |