Ahmad Yousef Alhilal

dblp:297/7202 · also Ahmad Alhilal · DBLP profile ↗
← Back
16ranked-venue papers
4as first author
16since 2021 · last 2026
0000-0002-2575-1391ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 2 first-author · 7 since 2021Human-computer interaction and ubiquitous computing · 5 · 5 since 2021Computer networks · 3 · 3 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Enhancing Creativity in Virtuality: How Annotations in Creative Support Tool Innovate Design Ideation in Virtual Reality
abstract
Asynchronous design ideation is increasingly vital in virtual reality (VR) environments. Annotations, referring to notes, marks, or comments overlaid on virtual designs, are essential for providing context, feedback, and clarity, facilitating communication among designers using Creativity Support Tools (CSTs) in the design sector. However, the effectiveness of various annotation methods in enhancing creativity during VR design ideation remains underexplored. This study aims to identify effective annotation methods that promote creativity in VR-based design ideation, addressing the shift from traditional paper-and-pen practices to immersive environments. We adopted a three-phase approach: (1) semi-structured interviews with design experts, (2) an empirical mixed-design study with design professionals, leading to the development of AsyncCreativity, our VR CST system, and (3) a deployment study evaluating annotations via AsyncCreativity in VR design ideation. Our findings reveal that multimodal annotations, especially audio annotations, significantly enhance user engagement and creativity compared to unimodal annotations. Participants reported improved ideation experiences when utilizing diverse annotation types during ideation when searching, sketching, and presenting. Our study provides valuable insights for developing effective CSTs in VR and design field, advancing the understanding of how optimized annotation strategies can foster creativity in immersive design environments.
Xuetong Wang, Ching Christie Pang, Ahmad Yousef Alhilal, Simin Yang, Tristan Braud, Pan Hui 0001
VR4
2026 Frame Complexity-Aware Foveated Video Encoding for Real-time High-Quality Streaming
abstract
VR streaming and VR cloud gaming require high-resolution video streaming to provide users with high quality visual experience and maximize their interaction and immersion. Consequently, the video streams have a high bitrate and require a large amount of available bandwidth. Foveated video encoding (FVE) reduces bandwidth demand by selectively allocating higher quality to perceptually relevant regions based on human visual characteristics. However, scenes with high spatial detail or rapid motion introduce spatial and temporal frame complexities. Conventional video encoding assesses complexity and manages quality and bitrate through an internal rate-control mechanism. However, the SoTA FVE methods perform the quality allocation process after the rate control has already run. This may lead to rate violations and, consequently, under-or overutilization of available bandwidth, which in turn causes increased latency and/or reduced visual quality. In this paper, we present a real-time complexity-adaptive FVE method that minimizes computational latency using a GPU-accelerated compute shader. By prioritizing key spatial and temporal complexity variables, our weighted complexity estimation optimizes quality assignment. Our method outperforms the complexity-agnostic FVE benchmark with 44-78% greater bitrate stability, 21% lower latency, and 7% higher perceptual quality. It also surpasses the partial complexity-aware FVE benchmark, delivering 18-94% better network utilization alongside a 7% latency reduction and a 4-5% gain in quality. Our method also ensures generalizability across diverse scenarios.
Ze Wu 0006, Ahmad Yousef Alhilal, Yuk Hang Tsui, Matti Siekkinen, Pan Hui 0001
VR2
2026 FovRL: Joint Foveation and Quality Control for Immersive VR Streaming Using Reinforcement Learning
abstract
VR cloud gaming promises immersive experiences, yet its realization is critically challenged by the trade-off between stringent latency requirements and high visual quality under unpredictable network conditions. Existing heuristic adaptive bitrate and foveation approaches lack adaptability to highly dynamic mobile networks. This results in a suboptimal trade-off between bandwidth usage and visual quality. While data-driven approaches (i.e., reinforcement learning, RL) have been successful in video streaming, their application to VR cloud gaming poses particular challenges. The stringent demands for high resolution and frame rate, and ultra-low latency are compounded by the necessity for fine-grained, per-frame inference to adapt to rapid changes in user gaze and network conditions. This work introduces FovRL, an RL framework for jointly optimizing foveation parameters and bitrate allocation in response to real-time network throughput. Our work pioneers the application of RL for real-time foveated encoding in immersive VR cloud gaming. Evaluations over real-world networks reveal that FovRL enhances bitrate adaptability to deliver superior perceptual visual quality, while maintaining latency comparable to the SoTA.
Yuk Hang Tsui, Ze Wu 0006, Ahmad Yousef Alhilal, Matti Siekkinen, Pan Hui 0001
WWW3
2026 Saliency-Guided Foveated Video Encoding for Low-Latency and Immersive Cloud VR
abstract
In cloud virtual reality (VR), delivering high perceptual visual quality under constrained wireless bandwidth remains a pivotal challenge. Foveated video encoding (FVE) copes with this by human vision-driven quality allocation, reducing bandwidth demand while preserving high visual fidelity for immersive VR streaming. However, existing FVE methods are content-independent, relying exclusively on gaze position and predefined quality degradation profiles. Thus, they fall short in modeling the intricate mechanisms of human attention. This leads to suboptimal quality distribution for visual saliency stimuli, causing perceptible artifacts and degraded perceptual quality. In this paper, we introduce a novel AI-driven streaming framework that incorporates visual saliency cues, defined as the innate capacity of scene elements to attract attention, into the video encoding pipeline. To meet the stringent low-latency demands of immersive VR, we propose a lightweight deep neural network for saliency inference, reducing computational complexity (FLOPs) by 48× compared to prior models while maintaining comparable accuracy. We integrate our pipeline into an open-source cloud VR gaming platform and conduct comprehensive experiments. Evaluation results demonstrate that our approach enhances perceptual visual quality by 22.98% compared to SoTA systems. Our IRB-approved user study shows that the saliency-guided FVE pipeline achieves superior visual quality and spatial smoothness while significantly reducing noticeable artifacts introduced by gaze-exclusive FVE. The project source code is available at https://github.com/WuZemyp/SaliencyFov.
Ze Wu 0006, Ahmad Yousef Alhilal, Yuk Hang Tsui, Wen Jye Chai, Matti Siekkinen, Pan Hui 0001
IEEE Trans. Vis. Comput. Graph.2
2025 Congestion Control for VR Cloud Gaming: Integration and Comparison in Real VR Gaming Environment
abstract
Virtual reality (VR) cloud gaming is increasingly developing in the gaming industry. Yet, the performance of the congestion control algorithms on top of which these systems build remains under-explored. In this study, we implement two industry-standard network congestion control algorithms, Google Congestion Control (GCC) and Network-Assisted Dynamic Adaptation (NADA), according to their Requests for Comments (RFCs), and integrate them into an open-source VR gaming system (ALVR). Including ALVR's congestion control (ALVR-ABR), we conduct extensive experiments on real-world networks to evaluate each algorithm's frame latency, target-to-receiving bitrate gap, dropped frames, image quality, and fairness among heterogeneous competing flows. GCC decreases frame latency by 352ABR present significant gaps between the selected and received bitrate, causing substantial congestion-induced frame drops, while GCC has a minimal gap, resulting in minor frame drops, suggesting its suitability for game-player interaction. GCC exhibits a 2.7ABR, respectively, indicating slight immersion degradation. However, only NADA ensures a fair bandwidth share against loss-based flows due to its bitrate response to loss-induced congestion signals and lower sensitivity to delay gradients compared to GCC.
Ahmad Yousef Alhilal, Ze Wu 0006, Teemu Kämäräinen, Tristan Braud, Matti Siekkinen
ACM Multimedia1
2025 Gaze-Adaptive Foveation for Remote Rendered VR
abstract
Remote rendering enables high-fidelity virtual reality (VR) experiences on standalone headsets by offloading intensive graphics workloads to remote servers. However, streaming high-quality VR graphics imposes substantial bandwidth and latency challenges. Spatial compression is a form of foveation which addresses this challenge by leveraging the human visual system's varying acuity, allocating higher visual quality around the user's gaze while reducing resolution in the periphery. In this work, we implement three gaze-adaptive foveation methods: Dynamic Axis-Aligned Distortion Transmission (D-AADT2 and D-AADT3) and Dynamic Foveated Radial Warp (D-FRW)) of which only D-AADT2 has been previously presented. These methods dynamically adapt spatial compression based on gaze-tracking input, ensuring optimal perceptual quality. We integrate these methods together with their static counterparts into the open-source Air Light VR (ALVR) remote-rendering framework, enabling native (72 FPS) framerates. We conclude a comprehensive objective evaluation across diverse VR games and demonstrate that the dynamic methods significantly outperform traditional static approaches in both encoding efficiency and perceptual quality metrics. A complementary subjective user study further validates these findings, confirming that dynamic gaze-adaptive foveation substantially enhances visual quality, immersion, and user interaction experience.
Adhi Widagdo, Teemu Kämäräinen, Ahmad Yousef Alhilal, Matti Siekkinen, Cheng-Hsin Hsu
ACM Multimedia3
2024 FovOptix: Human Vision-Compatible Video Encoding and Adaptive Streaming in VR Cloud Gaming
abstract
VR cloud gaming enables users to play high-end VR games on lightweight devices by offloading rendering tasks to cloud servers. Despite video compression, high-definition video streaming requires substantial data transfer rates. Foveated rendering (FR) and video encoding (FVE) leverage the non-uniform perception of the human visual system to reduce computing and bandwidth demand. They enhance visual quality in central gaze regions and reduce it in the periphery. However, bandwidth variation may hinder the provision of smooth VR gaming experiences. We present FovOptix, a system that combines FR with adaptive FVE to deliver video stream at a lower yet adaptive bitrate while not compromising the perceived video quality. FovOptix is based on a game-agnostic open-source to ensure reproducibility and compatibility with various games. We evaluate FovOptix against benchmarks using 5G mobile network traces. FovOptix achieves a latency reduction of 3% compared to the Google standard and a significant +100% reduction compared to other solutions. Additionally, it enhances the visual quality within the player's region of interest. Consequently, FovOptix attains the highest playability and gaming scores while minimizing the severity of motion sickness. FovOptix thus offers smooth and accessible VR cloud gaming for a wider range of players.
Ahmad Yousef Alhilal, Ze Wu 0006, Yuk Hang Tsui, Pan Hui 0001
MMSys1
2024 AnchorLoc: Large-Scale, Real-Time Visual Localisation Through Anchor Extraction and Detection
abstract
Pervasive Augmented Reality (AR) requires accurate pose registration of the device in real-time at a neighbourhood-to-city scale. At such a scale, most pose registration techniques suffer from exponential computational and storage costs and a significant data collection burden. This paper introduces AnchorLoc, a framework that relies on visual anchors (stable and highly recognisable visual elements in a scene) to perform fast and accurate pose registration. Anchorloc automatically identifies these anchors from large image sequences to optimise the search space in later image retrieval and pose registration. As such, it significantly improves the computational efficiency of existing hierarchical localisation pipelines without compromising accuracy. We collect a large-scale localisation dataset consisting of image sequences and 3D reconstruction of a university campus. AnchorLoc reduces localisation runtime by 83% on our campus dataset and 69% on the Cambridge Landmarks dataset without significantly increasing mean pose estimation errors. It is also more accurate and faster than SLD, a localisation algorithm that takes a comparable approach at the keypoint level. This work informs the development of more efficient pervasive AR applications that rely on both absolute and relative camera pose tracking on image sequences.
Chun Ho Park, Ahmad Yousef Alhilal, Tristan Braud, Pan Hui 0001
PerCom2
2024 The Jade Gateway to Exergaming: How Socio-Cultural Factors Shape Exergaming Among East Asian Older Adults
abstract
Exergaming, blending exercise and gaming, improves the physical and mental health of older adults. We currently do not fully know the factors that drive older adults to either engage in or abstain from exergaming. Large-scale studies investigating this are still scarce, particularly those studying East Asian older adults. To address this, we interviewed 64 older adults from China, Japan, and South Korea about their attitudes toward exergames. Most participants viewed exergames with a positive inquisitiveness. However, socio-cultural factors can obstruct this curiosity. Our study shows that perceptions of aging, lifestyle, the presence of support networks, and the cultural relevance of game mechanics are the crucial factors influencing their exergame engagement. Thus, we stress the value of socio-cultural sensitivity in game design and urge the HCI community to adopt more diverse design practices. We provide several design suggestions for creating more culturally approachable exergames.
Reza Hadi Mogavi, Juhyung Son, Simin Yang, Derrick M. Wang, Lydia Choong, Ahmad Yousef Alhilal, Peng Yuan Zhou, Pan Hui 0001, Lennart E. Nacke
Proc. ACM Hum. Comput. Interact.6
2024 Attention-Based QoE-Aware Digital Twin Empowered Edge Computing for Immersive Virtual Reality
abstract
Metaverse applications such as virtual reality (VR) content streaming, require optimal resource allocation strategies for mobile edge computing (MEC) to ensure a high-quality user experience. In contrast to online reinforcement learning (RL) algorithms, which can incur substantial communication overheads and longer delays, the majority of existing works employ offline-trained RL algorithms for resource allocation decisions in MEC systems. However, they neglect the impact of desynchronization between the physical and digital worlds on the effectiveness of the allocation strategy. In this paper, we tackle this desynchronization using a continual RL (CRL) framework that facilitates the resource allocation dynamically for MEC-enabled VR content streaming. We first design a digital twin-empowered edge computing (DTEC) system and formulate a quality of experience (QoE) maximization problem based on attention-based resolution perception. This problem optimizes the allocation of computing and bandwidth resources while adapting the attention-based resolution of the VR content. The CRL framework in DTEC enables adaptive online execution in a time-varying environment. We propose three variants of CRL, namely Continual Deep Deterministic Policy Gradient (CDDPG), Prioritized Experience Replay - CDDPG (PER-CDDPG), and Freshness Prioritized Experience Replay - CDDPG (FPER-CDDPG). We evaluate these algorithms, including two other benchmarks, using extensive experiments. FPER-CDDPG shows superior performance in terms of average latency, QoE, and successful delivery rate as well as meeting the hfQoE requirements over long-term execution while ensuring system scalability with the increasing number of users.
Jiadong Yu, Ahmad Yousef Alhilal, Tailin Zhou, Pan Hui 0001, Danny H. K. Tsang
IEEE Trans. Wirel. Commun.2
2023 QoE Optimization for VR Streaming: a Continual RL Framework in Digital Twin-empowered MEC
abstract
Mobile edge computing (MEC) resource allocation for remote rendering in virtual reality (VR) content streaming is critical for user experience. However, resource allocation becomes challenging due to the desynchronization between the physical and digital worlds in digital twin-empowered MEC. This paper presents our continual RL framework that facilitates dynamic resource allocation for MEC-enabled VR content streaming. We first design a digital twin-empowered edge computing (DTEC) system and formulate a maximization problem that considers attention-based resolution perception to maximize the quality of experience (QoE). This problem optimizes the allocation of computing and bandwidth resources while adapting the attention-based resolution of the VR content. We then apply continual reinforcement learning (CRL) to enable adaptive attention-based resolution VR streaming in a time-varying environment. We base the CRL's reward function on the QoE and horizon-fairness QoE (hfQoE) constraints. We support CRL with prioritized experience replay - continual deep deterministic policy gradient (PER-CDDPG) to enhance the performance of continual learning in the presence of time-varying DT updates. We test PER-CDDPG using extensive experiments and evaluation. PER-CDDPG outperforms the benchmarks in terms of average latency, QoE, and successful delivery rate as well as meeting the hfQoE requirements and performance over long-term execution while ensuring system scalability with the increasing number of users.
Jiadong Yu, Ahmad Yousef Alhilal, Tailin Zhou, Pan Hui 0001, Danny H. K. Tsang
GLOBECOM2
2022 Beyond the Blue Sky of Multimodal Interaction: A Centennial Vision of Interplanetary Virtual Spaces in Turn-based Metaverse
abstract
Human habitation across multiple planets requires communication and social connection between planets. When the infrastructure of a deep space network becomes mature, immersive cyberspace, known as the Metaverse, can exchange diversified user data and host multitudinous virtual worlds. Nevertheless, such immersive cyberspace unavoidably encounters latency in minutes, and thus operates in a turn-taking manner. This Blue Sky paper illustrates a vision of an interplanetary Metaverse that connects Earthian and Martian users in a turn-based Metaverse. Accordingly, we briefly discuss several grand challenges to catalyze research initiatives for the ‘Digital Big Bang’ on Mars.
Lik-Hang Lee, Carlos Bermejo 0001, Ahmad Yousef Alhilal, Tristan Braud, Simo Hosio, Esmée Henrieke Anne de Haas, Pan Hui 0001
ICMI3
2022 Human-Avatar Interaction in Metaverse: Framework for Full-Body Interaction
abstract
The metaverse is a network of shared virtual environments where people can interact synchronously through their avatars. To enable this, it is necessary to accurately capture and recreate (physical) human motion. This is used to render avatars correctly, reflecting the motion of their corresponding users. In large-scale environments this must be done in real-time. This paper proposes a human-avatar framework with full-body motion capture. Its goal is to deliver high-accuracy capture with low computational and network overheads. It relies on a lightweight Octree data structure to record and transmit motion to other users. We conduct a user study with 22 participants and perform a preliminary evaluation of its scalability. Our user study shows that Octree with Inverse Kinematic achieves the best trade-off, achieving low delay and high accuracy. Our proposed solution delivers the lowest delay, with an average of 67ms in an environment of 8 concurrent users. It attains a 55.7% improvement over the prior techniques.
Kit-Yung Lam, Ahmad Yousef Alhilal, Lik-Hang Lee, Gareth Tyson, Pan Hui 0001
MMAsia3
2022 Nebula: Reliable Low-latency Video Transmission for Mobile Cloud Gaming
abstract
Mobile cloud gaming enables high-end games on constrained devices by streaming the game content from powerful servers through mobile networks. Mobile networks suffer from highly variable bandwidth, latency, and losses that affect the gaming experience. This paper introduces , an end-to-end cloud gaming framework to minimize the impact of network conditions on the user experience. relies on an end-to-end distortion model adapting the video source rate and the amount of frame-level redundancy based on the measured network conditions. As a result, it minimizes the motion-to-photon (MTP) latency while protecting the frames from losses. We fully implement and evaluate its performance against the state-of-the-art techniques and latest research in real-time mobile cloud gaming transmission on a physical testbed over emulated and real wireless networks. consistently balances MTP latency (<140 ms) and visual quality (>31dB) even in highly variable environments. A user experiment confirms that maximizes the user experience with high perceived video quality, playability, and low user load.
Ahmad Yousef Alhilal, Tristan Braud, Bo Han 0001, Pan Hui 0001
WWW1
2021 Talaria: in-engine synchronisation for seamless migration of mobile edge gaming instances
abstract
Mobile cloud gaming requires a very low end-to-end latency. Edge computing significantly reduces network latency. However, in mobility scenarios, the user will frequently move out of the edge server's coverage area, requiring frequent migration of the game instance. This paper presents Talaria, an in-engine content synchronisation solution for unnoticeable game instance migration between edge servers. Talaria creates a minimal instance with content immediately relevant to the game experience, allowing the client to switch servers in a minimal amount of time. The remaining content is then synchronised according to priority until the game's state is coherent between both instances. Our implementation of Talaria as a Unity engine plugin reduces the game's downtime by 61% compared to one-off server migration, with an average latency below 25 ms for the server migration, and 87 ms for the entire game synchronisation.
Tristan Braud, Ahmad Yousef Alhilal, Pan Hui 0001
CoNEXT2
2021 CAD3: Edge-facilitated Real-time Collaborative Abnormal Driving Distributed Detection
abstract
Speeding, slowing down, and sudden acceleration are the leading causes of fatal accidents on highways. Anomalous driving behavior detection can improve road safety by informing drivers who are in the vicinity of dangerous vehicles. However, detecting abnormal driving behavior at the city-scale in a centralized fashion results in considerable network and computation load, that would significantly restrict the scalability of the system. In this paper, we propose CAD3, a distributed collaborative system for road-aware and driver-aware anomaly driving detection. CAD3 considers a decentralized deployment of edge computation nodes on the roadside and combines collaborative and context-aware computation with low-latency communication to detect and inform nearby drivers of unsafe behaviors of other vehicles in real-time. Adjacent edge nodes collaborate to improve the detection of abnormal driving behavior at the city-scale. We evaluate CAD3 with a physical testbed implementation. We emulate realistic driving scenarios from a real driving data set of 3,000 vehicles, 214,000 trips, and 18 million trajectories of private cars in Shenzhen, China. At the microscopic (road) level, CAD3 significantly improves the accuracy of detection and lowers the number of potential accidents caused by false negatives up to four times and 24 times as compared to distributed standalone and centralized models, respectively. CAD3 can scale up to 256 vehicles connected to a single node while keeping the end-to-end latency under 50 ms and a required bandwidth below 5 mbps. At the mesoscopic (driver-trip) level, CAD3 performs stable and accurate detection over time, owing to local RSU interaction. With a dense deployment of edge nodes, CAD3 can scale up to the size of Shenzhen, a megalopolis of 12 million inhabitant with over 2 million concurrent vehicles at peak hours.
Ahmad Yousef Alhilal, Tristan Braud, Xiang Su 0001, Luay Al Asadi, Pan Hui 0001
ICDCS1