Teemu Kämäräinen

dblp:157/0335 · DBLP profile ↗
← Back
16ranked-venue papers
2as first author
7since 2021 · last 2025
0000-0003-4685-6763ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 11 · 2 first-author · 6 since 2021Computer networks · 8 · 3 since 2021
YearPublicationVenuePosition
2025 Congestion Control for VR Cloud Gaming: Integration and Comparison in Real VR Gaming Environment
abstract
Virtual reality (VR) cloud gaming is increasingly developing in the gaming industry. Yet, the performance of the congestion control algorithms on top of which these systems build remains under-explored. In this study, we implement two industry-standard network congestion control algorithms, Google Congestion Control (GCC) and Network-Assisted Dynamic Adaptation (NADA), according to their Requests for Comments (RFCs), and integrate them into an open-source VR gaming system (ALVR). Including ALVR's congestion control (ALVR-ABR), we conduct extensive experiments on real-world networks to evaluate each algorithm's frame latency, target-to-receiving bitrate gap, dropped frames, image quality, and fairness among heterogeneous competing flows. GCC decreases frame latency by 352ABR present significant gaps between the selected and received bitrate, causing substantial congestion-induced frame drops, while GCC has a minimal gap, resulting in minor frame drops, suggesting its suitability for game-player interaction. GCC exhibits a 2.7ABR, respectively, indicating slight immersion degradation. However, only NADA ensures a fair bandwidth share against loss-based flows due to its bitrate response to loss-induced congestion signals and lower sensitivity to delay gradients compared to GCC.
Ahmad Yousef Alhilal, Ze Wu 0006, Teemu Kämäräinen, Tristan Braud, Matti Siekkinen
ACM Multimedia3
2025 Gaze-Adaptive Foveation for Remote Rendered VR
abstract
Remote rendering enables high-fidelity virtual reality (VR) experiences on standalone headsets by offloading intensive graphics workloads to remote servers. However, streaming high-quality VR graphics imposes substantial bandwidth and latency challenges. Spatial compression is a form of foveation which addresses this challenge by leveraging the human visual system's varying acuity, allocating higher visual quality around the user's gaze while reducing resolution in the periphery. In this work, we implement three gaze-adaptive foveation methods: Dynamic Axis-Aligned Distortion Transmission (D-AADT2 and D-AADT3) and Dynamic Foveated Radial Warp (D-FRW)) of which only D-AADT2 has been previously presented. These methods dynamically adapt spatial compression based on gaze-tracking input, ensuring optimal perceptual quality. We integrate these methods together with their static counterparts into the open-source Air Light VR (ALVR) remote-rendering framework, enabling native (72 FPS) framerates. We conclude a comprehensive objective evaluation across diverse VR games and demonstrate that the dynamic methods significantly outperform traditional static approaches in both encoding efficiency and perceptual quality metrics. A complementary subjective user study further validates these findings, confirming that dynamic gaze-adaptive foveation substantially enhances visual quality, immersion, and user interaction experience.
Adhi Widagdo, Teemu Kämäräinen, Ahmad Yousef Alhilal, Matti Siekkinen, Cheng-Hsin Hsu
ACM Multimedia2
2023 Learning to Predict Head Pose in Remotely-Rendered Virtual Reality
abstract
Accurate characterization of Head Mounted Display (HMD) pose in a virtual scene is essential for rendering immersive graphics in Extended Reality (XR). Remote rendering employs servers in the cloud or at the edge of the network to overcome the computational limitations of either standalone or tethered HMDs. Unfortunately, it increases the latency experienced by the user; for this reason, predicting HMD pose in advance is highly beneficial, as long as it achieves high accuracy. This work provides a thorough characterization of solutions that forecast HMD pose in remotely-rendered virtual reality (VR) by considering six degrees of freedom. Specifically, it provides an extensive evaluation of pose representations, forecasting methods, machine learning models, and the use of multiple modalities along with joint and separate training. In particular, a novel three-point representation of pose is introduced together with a data fusion scheme for long-term short-term memory (LSTM) neural networks. Our findings show that machine learning models benefit from using multiple modalities, even though simple statistical models perform surprisingly well. Moreover, joint training is comparable to separate training with carefully chosen pose representation and data fusion strategies.
Gazi Karam Illahi, Ashutosh Vaishnav, Teemu Kämäräinen, Matti Siekkinen, Mario Di Francesco
MMSys3
2023 Will Dynamic Foveation Boost Cloud VR Gaming Experience?
abstract
Cloud Virtual Reality (VR) gaming offloads the computationally-intensive rendering tasks from resource-limited Head-Mounted Displays (HMDs) to cloud servers, which consume a staggering amount of bandwidth for high-quality gaming experiences. One way to cope with such high bandwidth demands is to capitalize on human vision systems by allocating a higher bitrate to the foveal region of HMD viewport, which is known as foveation in the literature. Although foveation was employed by remote VR gaming, existing open-source projects all adopt static foveation, in which the HMD gamer gaze position is assumed to be fixed at the viewport center. In this paper, we construct the very first cloud VR gaming system that supports dynamic foveation. That is, the real-time gaze positions of gamers are streamed from eye-trackers on HMDs to cloud servers, which in turn adjust the foveation parameters, such as foveal region size/location and peripheral region quality degradation, accordingly. Using our developed cloud VR gaming system, we design and carry out a user study using a game called Fruit Ninja VR 2 to find the foveation parameters in static and dynamic foveation for maximizing the gaming Quality of Experience (QoE) in Mean Opinion Score (MOS). With the chosen foveation parameters, we found that, compared to cloud VR gaming without foveation, static foveation leads to a MOS increase of 0.60 and a bitrate reduction of 8.71%. Furthermore, adopting dynamic foveation results in an additional 0.60 increase on MOS while saving 9.81% bitrate, compared to static foveation. Our findings demonstrate the potential of dynamic foveation in cloud VR gaming, which dictates both high visual quality and short response time. The optimization techniques developed in this and follow-up work could benefit other cloud-rendered applications that typically have less strict requirements than cloud VR gaming.
Eric Jia-Wei Fang, Kuan-Yu Lee, Teemu Kämäräinen, Matti Siekkinen, Cheng-Hsin Hsu
NOSSDAV3
2023 Neural Network Assisted Depth Map Packing for Compression Using Standard Hardware Video Codecs
abstract
Depth maps are needed by various graphics rendering and processing operations. Depth map streaming is often necessary when such operations are performed in a distributed system and it requires in most cases fast performing compression, which is why video codecs are often used. Hardware implementations of standard video codecs enable relatively high resolution and frame rate combinations, even on resource constrained devices, but unfortunately those implementations do not currently support RGB+depth extensions. However, they can be used for depth compression by first packing the depth maps into RGB or YUV frames. We investigate depth map compression using a combination of depth map packing followed by encoding with a standard video codec. We show that the precision at which depth maps are packed has a large and nontrivial impact on the resulting error caused by the combination of the packing scheme and lossy compression when the bitrate is constrained. Consequently, we propose a variable precision packing scheme assisted by a neural network model that predicts the optimal precision for each depth map given a bitrate constraint. We demonstrate that the model yields near optimal predictions and that it can be integrated into a game engine with very low overhead using modern hardware.
Matti Siekkinen, Teemu Kämäräinen
ACM Trans. Multim. Comput. Commun. Appl.2
2021 Foveated streaming of real-time graphics
abstract
Remote rendering systems comprise powerful servers that render graphics on behalf of low-end client devices and stream the graphics as compressed video, enabling high end gaming and Virtual Reality on those devices. One key challenge with them is the amount of bandwidth required for streaming high quality video. Humans have spatially non-uniform visual acuity: We have sharp central vision but our ability to discern details rapidly decreases with angular distance from the point of gaze. This phenomenon called foveation can be taken advantage of to reduce the need for bandwidth. In this paper, we study three different methods to produce a foveated video stream of real-time rendered graphics in a remote rendered system: 1) foveated shading as part of the rendering pipeline, 2) foveation as post processing step after rendering and before video encoding, 3) foveated video encoding. We report results from a number of experiments with these methods. They suggest that foveated rendering alone does not help save bandwidth. Instead, the two other methods decrease the resulting video bitrate significantly but they also have different quality per bit and latency profiles, which makes them desirable solutions in slightly different situations.
Gazi Karam Illahi, Matti Siekkinen, Teemu Kämäräinen, Antti Ylä-Jääski
MMSys3
2021 Multi-Tier CloudVR: Leveraging Edge Computing in Remote Rendered Virtual Reality
abstract
The availability of high bandwidth with low-latency communication in 5G mobile networks enables remote rendered real-time virtual reality (VR) applications. Remote rendering of VR graphics in a cloud removes the need for local personal computer for graphics rendering and augments weak graphics processing unit capacity of stand-alone VR headsets. However, to prevent the added network latency of remote rendering from ruining user experience, rendering a locally navigable viewport that is larger than the field of view of the HMD is necessary. The size of the viewport required depends on latency: Longer latency requires rendering a larger viewport and streaming more content. In this article, we aim to utilize multi-access edge computing to assist the backend cloud in such remote rendered interactive VR. Given the dependency between latency and amount and quality of the content streamed, our objective is to jointly optimize the tradeoff between average video quality and delivery latency. Formulating the problem as mixed integer nonlinear programming, we leverage the interpolation between client’s field of view frame size and overall latency to convert the problem to integer nonlinear programming model and then design efficient online algorithms to solve it. The results of our simulations supplemented by real-world user data reveal that enabling a desired balance between video quality and latency, our algorithm particularly achieves the improvements of on average about 22% and 12% in term of video delivery latency and 8% in term of video quality compared to respectively order-of-arrival, threshold-based, and random-location strategies.
Abbas Mehrabi, Matti Siekkinen, Teemu Kämäräinen, Antti Ylä-Jääski
ACM Trans. Multim. Comput. Commun. Appl.3
2020 On the Interplay of Foveated Rendering and Video Encoding
abstract
Humans have sharp central vision but low peripheral visual acuity. Prior work has taken advantage of this phenomenon in two ways: foveated rendering (FR) reduces the computational workload of rendering by producing lower visual quality for peripheral regions and foveated video encoding (FVE) reduces the bitrate of streamed video through heavier compression of peripheral regions. Remote rendering systems require both rendering and video encoding and the two techniques can be combined to reduce both computing and bandwidth consumption. We report early results from such a combination with remote VR rendering. The results highlight that FR causes large bitrate overhead when combined with normal video encoding but combining it with FVE can mitigate it.
Gazi Karam Illahi, Matti Siekkinen, Teemu Kämäräinen, Antti Ylä-Jääski
VRST3
2019 Multi-carrier Measurement Study of Mobile Network Latency: The Tale of Hong Kong and Helsinki
abstract
Real time interactive cloud-based mobile applications such as augmented reality and cloud gaming require low and stable latency, especially in urban areas. These conditions are difficult to meet with the traditional single carrier LTE network access and consolidated server deployment in a cloud. Yet, with multiple SIM/multiple radio devices, latency can be kept under a given threshold through dynamic selection among multiple carriers and server deployment at network edge. To this end, it is necessary to understand how mobile network latency changes over time during a session with different carriers and how the server placement affects the latencies. In this paper, we present results from a measurement study of mobile network latency and jitter in 4G networks of Hong Kong and Helsinki, two very different cities in terms of population density and mobile infrastructure. Based on the results, we introduce a lightweight carrier selection algorithm that displays latencies 10 to 20% lower than single carrier operation.
Tristan Braud, Teemu Kämäräinen, Matti Siekkinen, Pan Hui 0001
MSN2
2018 CloudVR: Cloud Accelerated Interactive Mobile Virtual Reality
abstract
High quality immersive Virtual Reality experience currently requires a PC setup with cable connected head mounted display, which is expensive and restricts user mobility. This paper presents CloudVR which is a system for cloud accelerated interactive mobile VR. It is designed to provide short rotation and interaction latencies through panoramic rendering and dynamic object placement. CloudVR also includes rendering optimizations to reduce server-side computational load and bandwidth requirements between the server and client. Performance measurements with a CloudVR prototype suggest that the optimizations make it possible to double the server's framerate and halve the amount of bandwidth required and that small objects can be quickly moved at run time to client device for rendering to provide shorter interaction latency. A small-scale user study indicates that CloudVR users do not notice small network latencies (20ms) and even much longer ones (100-200ms) become non-trivial to detect when they do not affect the interaction with objects. Finally, we present a design of CloudVR extension to multi-user scenarios.
Teemu Kämäräinen, Matti Siekkinen, Jukka Eerikäinen, Antti Ylä-Jääski
ACM Multimedia1
2018 Latency and throughput characterization of convolutional neural networks for mobile computer vision
abstract
We study performance characteristics of convolutional neural networks (CNN) for mobile computer vision systems. CNNs have proven to be a powerful and efficient approach to implement such systems. However, the system performance depends largely on the utilization of hardware accelerators, which are able to speed up the execution of the underlying mathematical operations tremendously through massive parallelism. Our contribution is performance characterization of multiple CNN-based models for object recognition and detection with several different hardware platforms and software frameworks, using both local (on-device) and remote (network-side server) computation. The measurements are conducted using real workloads and real processing platforms. On the platform side, we concentrate especially on TensorFlow and TensorRT. Our measurements include embedded processors found on mobile devices and high-performance processors that can be used on the network side of mobile systems. We show that there exists significant latency-throughput trade-offs but the behavior is very complex. We demonstrate and discuss several factors that affect the performance and yield this complex behavior.
Jussi Hanhirova, Teemu Kämäräinen, Sipi Seppälä, Matti Siekkinen, Vesa Hirvisalo, Antti Ylä-Jääski
MMSys2
2018 Can You See What I See? Quality-of-Experience Measurements of Mobile Live Video Broadcasting
abstract
Broadcasting live video directly from mobile devices is rapidly gaining popularity with applications like Periscope and Facebook Live. The quality of experience (QoE) provided by these services comprises many factors, such as quality of transmitted video, video playback stalling, end-to-end latency, and impact on battery life, and they are not yet well understood. In this article, we examine mainly the Periscope service through a comprehensive measurement study and compare it in some aspects to Facebook Live. We shed light on the usage of Periscope through analysis of crawled data and then investigate the aforementioned QoE factors through statistical analyses as well as controlled small-scale measurements using a couple of different smartphones and both versions, Android and iOS, of the two applications. We report a number of findings including the discrepancy in latency between the two most commonly used protocols, RTMP and HLS, surprising surges in bandwidth demand caused by the Periscope app’s chat feature, substantial variations in video quality, poor adaptation of video bitrate to available upstream bandwidth at the video broadcaster side, and significant power consumption caused by the applications.
Matti Siekkinen, Teemu Kämäräinen, Leonardo Favario, Enrico Masala
ACM Trans. Multim. Comput. Commun. Appl.2
2017 A Measurement Study on Achieving Imperceptible Latency in Mobile Cloud Gaming
abstract
Cloud gaming is a relatively new paradigm in which the game is rendered in the cloud and is streamed to an end-user device through a thin client. Latency is a key challenge for cloud gaming. In order to optimize the end-to-end latency, it is first necessary to understand how the end-to-end latency builds up from the mobile device to the cloud gaming server. In this paper we dissect the delays occurring in the mobile device and measure access delays in various networks and network conditions. We also perform a Europe-wide latency measurement study to find the optimal server locations and see how the number of server locations affects the network delay. The results are compared to limits found for perceivable delays in recent human-computer interaction studies. We show that the limits can be achieved only with the latest mobile devices with specific control methods. In addition, we study the expected latency reduction by near future technological development and show that its potential impact is bigger on the end-to-end latency than that of replication of the service and server placement optimization.
Teemu Kämäräinen, Matti Siekkinen, Antti Ylä-Jääski, Pan Hui 0001
MMSys1
2017 Exploring Vision-Based Techniques for Outdoor Positioning Systems: A Feasibility Study
abstract
Recent advances from wearables have significantly changed the way how humans communicate with the surrounding environment. To some extent, they have extended and augmented the capability of humans. For example, with a Google Glass, people can take pictures simply by winking eyes twice, which releases human hands from the cumbersome image-taking process. Thus, it enables new application scenarios that were not possible before. In this paper, we investigate utilizing vision-based techniques to provide a wearable positioning system. Specifically, we propose a Human-centric Positioning System (HoPS) that utilizes traffic signposts together with context information for real-time positioning. Towards that direction, we make three primary contributions: (1) we make several important observations that guide our design of HoPS system; for example, we find out that approximately 40 percent of traffic signposts monopolize a cell tower, and there are at most six signposts within the coverage of a single cell tower; (2) we investigate the impact factors of object detection success rate, and find its correlation with image quality, and resolution; and (3) we design and implement HoPS and an advanced version of HoPS based on additional context information from Wi-Fi network, which we name HoPS-WiFi. Experimental results demonstrate the effectiveness of HoPS, especially HoPS-WiFi, which can estimate the relevant location correctly within 1.3 seconds.
Meina Song, Zhonghong Ou, Eduardo Castellanos, Tuomas Ylipiha, Teemu Kämäräinen, Matti Siekkinen, Antti Ylä-Jääski, Pan Hui 0001
IEEE Trans. Mob. Comput.5
2016 A First Look at Quality of Mobile Live Streaming Experience: the Case of Periscope
Matti Siekkinen, Enrico Masala, Teemu Kämäräinen
Internet Measurement Conference3
2015 Performance evaluation of remote display access for mobile cloud computing
Youming Lin, Teemu Kämäräinen, Mario Di Francesco, Antti Ylä-Jääski
Comput. Commun.2