EDBT 2026 Demo / reviewers in the wild / expert
Håkon Kvale Stensland
dblp:63/801
· DBLP profile ↗
40ranked-venue papers
4as first author
2since 2021 · last 2025
0000-0002-1085-8540ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 30 · 1 first-author · 1 since 2021Computer networks · 6 · 3 first-authorArtificial intelligence and machine learning · 4 · 1 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 since 2021Systems, architecture and hardware · 2 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
2 papers |
Distributed systems · 48% Memory systems · 37% Interconnection networks and networks-on-chip · 11% | |
| Computer graphics and multimedia
3 papers |
Computational photography and imaging · 48% Multimedia analysis and retrieval · 38% Image and video processing · 14% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Medical and health informatics · 100% | |
| Computer networks
1 paper |
Content delivery and video streaming · 77% Transport protocols and congestion control · 23% |
Topics — the 8 heaviest of 14, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Memory systems › shared memory
distributed shared memory |
0.4 | 1 | 2020 | SmartIO: Zero-overhead Device Sharing through PCIe Networking · ACM Trans. Comput. Syst. 2020 |
Medical and health informatics
gastrointestinal endoscopy |
0.4 | 1 | 2019 | ACM Multimedia BioMedia 2019 Grand Challenge Overview · ACM Multimedia 2019 |
Computational photography and imaging
high dynamic range imaging |
0.2 | 1 | 2014 | Real-Time HDR Panorama Video · ACM Multimedia 2014 |
Multimedia analysis and retrieval
object tracking |
0.2 | 1 | 2014 | Automatic Real-Time Zooming and Panning on Salient Objects from a Panoramic Video · ACM Multimedia 2014 |
Computational photography and imaging › image stitching
panoramic image stitching |
0.2 | 1 | 2014 | Real-Time HDR Panorama Video · ACM Multimedia 2014 |
Distributed systems
stream processing |
0.1 | 1 | 2011 | Processing of multimedia data using the P2G framework · ACM Multimedia 2011 |
Image and video processing
image enhancement |
0.1 | 1 | 2014 | Real-Time HDR Panorama Video · ACM Multimedia 2014 |
Cloud and datacenter computing
cluster resource management and scheduling |
0.0 | 1 | 2011 | Processing of multimedia data using the P2G framework · ACM Multimedia 2011 |
Methods — techniques the papers use, named apart from their topics
deep learning · 0.8non-transparent bridging · 0.4distributed shared memory · 0.4multi-camera panorama · 0.2automatic zooming and panning · 0.2HDR video algorithms · 0.2kernel language · 0.1dependency graph scheduling · 0.1proxy-based protocol translation · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Automatic Prompt Generation for Zero-Shot Single Object Frame Segmentation in Videos Using Classification Models: A Polyp Case StudyabstractVideo object segmentation is vital for applications like medical diagnostics, but acquiring dense pixel-level annotations, especially for specialized domains like polyp segmentation, remains a major bottleneck. Foundational models offer zero-shot segmentation but typically require manual prompting, which is impractical for long videos. We propose Map2VidSeg, a novel pipeline that automatically generates prompts from imagelevel classification labels. It leverages localization cues (attention maps/CAMs) from a trained image classifier (ViT/CNN) to create bounding box prompts. These guide an efficient model (YOLOE) with tracking (BOT-SORT) and bidirectional propagation for initial segmentation. Optionally, a high-fidelity model (SAM-2) refines these masks using temporal memory and fusion. Demonstrated on the challenging SUN-SEG benchmark, finetuned DINOv2 (ViT) prompts significantly outperform DenseNet-121 (CNN). Our best configuration (DINOv2+YOLOE+SAM-2 Bidirectional) achieves Dice/mIoU 0.76/0.70 (Easy Unseen) and 0.66/0.60 (Hard Unseen), showcasing the viability of robust video segmentation without segmentation training data. Hanna Borgli, Michael Riegler 0001, Håkon Kvale Stensland, Pål Halvorsen |
CBMS | 3 |
| 2025 | Better Image Segmentation with Classification: Guiding Zero-Shot Models Using Class Activation Maps
Hanna Borgli, Håkon Kvale Stensland, Pål Halvorsen |
MMM (5) | 2 |
| 2020 | Real-Time Detection of Events in Soccer Videos using 3D Convolutional Neural NetworksabstractIn this paper, we present an algorithm for automatically detecting events in soccer videos using 3D convolutional neural networks. The algorithm uses a sliding window approach to scan over a given video to detect events such as goals, yellow/red cards, and player substitutions. We test the method on three different datasets from SoccerNet, the Swedish Allsvenskan, and the Norwegian Eliteserien. Overall, the results show that we can detect events with high recall, low latency, and accurate time estimation. The trade-off is a slightly lower precision compared to the current state-of-the-art, which has higher latency and performs better when a less accurate time estimation can be accepted. In addition to the presented algorithm, we perform an extensive ablation study on how the different parts of the training pipeline affect the final results. Olav A. Norgård Rongved, Steven Alexander Hicks, Vajira Thambawita, Håkon Kvale Stensland, Evi Zouganeli, Dag Johansen, Michael Riegler 0001, Pål Halvorsen |
ISM | 4 |
| 2020 | PMData: a sports logging datasetabstractIn this paper, we present PMData: a dataset that combines traditional lifelogging data with sports-activity data. Our dataset enables the development of novel data analysis and machine-learning applications where, for instance, additional sports data is used to predict and analyze everyday developments, like a person's weight and sleep patterns; and applications where traditional lifelog data is used in a sports context to predict athletes' performance. PMData combines input from Fitbit Versa 2 smartwatch wristbands, the PMSys sports logging smartphone application, and Google forms. Logging data has been collected from 16 persons for five months. Our initial experiments show that novel analyses are possible, but there is still room for improvement. Vajira Thambawita, Steven Alexander Hicks, Hanna Borgli, Håkon Kvale Stensland, Debesh Jha, Martin Kristoffer Svensen, Svein Arne Pettersen, Dag Johansen, Håvard D. Johansen, Susann Dahl Pettersen, Simon Nordvang, Sigurd Pedersen, Anders T. Gjerdrum, Tor-Morten Grønli, Per Morten Fredriksen, Ragnhild Eg, Kjeld Hansen, Siri Fagernes, Christine Claudi, Andreas Biørn-Hansen, Duc-Tien Dang-Nguyen, Tomas Kupka, Hugo Hammer, Ramesh Jain 0001, Michael Riegler 0001, Pål Halvorsen |
MMSys | 4 |
| 2020 | SmartIO: Zero-overhead Device Sharing through PCIe NetworkingabstractThe large variety of compute-heavy and data-driven applications accelerate the need for a distributed I/O solution that enables cost-effective scaling of resources between networked hosts. For example, in a cluster system, different machines may have various devices available at different times, but moving workloads to remote units over the network is often costly and introduces large overheads compared to accessing local resources. To facilitate I/O disaggregation and device sharing among hosts connected using Peripheral Component Interconnect Express (PCIe) non-transparent bridges, we present SmartIO. NVMes, GPUs, network adapters, or any other standard PCIe device may be borrowed and accessed directly, as if they were local to the remote machines. We provide capabilities beyond existing disaggregation solutions by combining traditional I/O with distributed shared-memory functionality, allowing devices to become part of the same global address space as cluster applications. Software is entirely removed from the data path, and simultaneous sharing of a device among application processes running on remote hosts is enabled. Our experimental results show that I/O devices can be shared with remote hosts, achieving native PCIe performance. Thus, compared to existing device distribution mechanisms, SmartIO provides more efficient, low-cost resource sharing, increasing the overall system performance. Jonas Markussen, Lars Bjørlykke Kristiansen, Pål Halvorsen, Halvor Kielland-Gyrud, Håkon Kvale Stensland, Carsten Griwodz |
ACM Trans. Comput. Syst. | 5 |
| 2019 | Saga: An Open Source Platform for Training Machine Learning Models and Community-driven Sharing of TechniquesabstractWith the increasing popularity of machine learning in areas such as multimedia indexing and social media analysis, comes an increasing number of tools for developing and training the machine learning models. These tools assist in the selection and optimization of hyperparameters, but the user becomes locked into the platform's built-in solutions. Therefore, this demo presents an open-source, community-driven platform for sharing machine learning models, training techniques, and datasets. The platform, called Saga, provides a machine learning training pipeline such as those found in machine learning services provided by cloud providers. However, Saga also provides a store where users can upload and share their machine learning training methods. Additionally, the store allows users to rate and comment on other users' uploaded methods, as well as download and run them without any additional setup. In our demo, we will run a scenario where a user wants to train an image classifier using Saga. The demonstration will involve downloading methods from the store and displaying the pipeline provided by the Saga platform. Rune Johan Borgli, Håkon Kvale Stensland, Pål Halvorsen, Michael Riegler 0001 |
CBMI | 2 |
| 2019 | Semantic Analysis of Soccer News for Automatic Game Event ClassificationabstractWe are today overwhelmed with information, of which an important part is news. Sports news, in particular, has become very popular, where soccer makes up a big part of this coverage. For sports fans, it can be a time consuming and tedious to keep up with the news that they really care about. In this paper, we present different machine learning methods applied to soccer news from a Norwegian newspaper and a TV station's news site to summarize the content in a short and digestible manner. We present a system to collect, index, label, analyze, and present the collected news articles based on the content. We perform a thorough comparison between deep learning and traditional machine learning algorithms on text classification. Furthermore, we present a dataset of soccer news which was collected from two different Norwegian news sites and shared online. Aanund Jupskås Nordskog, Pål Halvorsen, Steven Alexander Hicks, Håkon Kvale Stensland, Hugo Hammer, Dag Johansen, Michael Riegler 0001 |
CBMI | 4 |
| 2019 | Performance of Data Enhancements and Training Optimization for Neural Network: A Polyp Detection Case StudyabstractDeep learning using neural networks is becoming more and more popular. It is frequently used in areas like video analysis, image retrieval, traffic forecast and speech recognition. In this respect, the learning and training process usually requires a lot of data. However, in many areas, data is scarce which is definitely the case in our medical application scenario, i.e., polyp detection in the gastrointestinal tract. Here, colorectal cancer is on the list of most common cancer types, and often, the cancer arises from benign, adenomatous polyps containing dysplastic cells. Detection and removal of polyps can therefore prevent the development of cancer. % Due to high cost, time consumption, patient discomfort and in-accuracy of existing procedures, researchers have started to explore systems for automatic polyp detection to assist and automate current examination procedures. Following the current gained traction for neural networks, and the typical lack of medical data, we explore how data enhancements affect the training and evaluation of the networks in terms of polyp detection accuracy and particularly if it can be used to increase the detection rate. We also experiment with how various training techniques can be used to increase performance. Our experimental results show how data enhancement and training optimization can be used to increase different aspects of the performance, but we also point out mechanisms that have no, and even a negative, effect. Fredrik Lund Henriksen, Rune Jensen, Håkon Kvale Stensland, Dag Johansen, Michael Riegler 0001, Pål Halvorsen |
CBMS | 3 |
| 2019 | ACM Multimedia BioMedia 2019 Grand Challenge OverviewabstractThe BioMedia 2019 ACM Multimedia Grand Challenge is the first in a series of competitions focusing on the use of multimedia for different medical use-cases. In this year's challenge, the participants are asked to develop efficient algorithms which automatically detect a variety of findings commonly identified in the gastrointestinal (GI) tract (a part of the human digestive system). The purpose of this task is to develop methods to aid medical doctors performing routine endoscopy inspections of the GI tract. In this paper, we give a detailed description of the four different tasks of this year's challenge, present the datasets used for training and testing, and discuss how each submission is evaluated both qualitatively and quantitatively. Steven Alexander Hicks, Michael Riegler 0001, Pia H. Smedsrud, Trine B. Haugen, Kristin Ranheim Randel, Konstantin Pogorelov, Håkon Kvale Stensland, Duc-Tien Dang-Nguyen, Mathias Lux, Andreas Petlund, Thomas de Lange, Peter Thelin Schmidt, Pål Halvorsen |
ACM Multimedia | 7 |
| 2018 | Autonomic Adaptation of Multimedia Content Adhering to Application Mobility
Francisco Javier Velázquez-García, Pål Halvorsen, Håkon Kvale Stensland, Frank Eliassen |
DAIS | 3 |
| 2018 | Dynamic Adaptation of Multimedia Presentations for Videoconferencing in Application MobilityabstractApplication mobility is the paradigm where users can move their running applications to heterogeneous devices in a seamless manner. This mobility involves dynamic context changes of hardware, network resources, user environment, and user preferences. In order to continue multimedia processing under these context changes, applications need to adapt not only the collection of media streams, i.e., multimedia presentation, but also their internal configuration to work on different hardware. We present the performance analysis to adapt a videoconferencing prototype application in a proposed adaptation control loop to autonomously adapt multimedia pipelines. Results show that the time spent to create an adaptation plan and execute it is in the order of hundreds of milliseconds. The reconfiguration of pipelines, compared to building them from scratch, is approximately 1000 times faster when re-utilizing already instantiated hardware-dependent components. Therefore, we conclude that the adaptation of multimedia pipelines is a feasible approach for multimedia applications that adhere to application mobility. Francisco Javier Velázquez-García, Pål Halvorsen, Håkon Kvale Stensland, Frank Eliassen |
ICME | 3 |
| 2018 | Tradeoffs Using Binary and Multiclass Neural Network Classification for Medical Multidisease DetectionabstractThe interest in neural networks has increased significantly, and the application of this type of machine learning is vast, ranging from natural image classification to medical image segmentation. However, many users of neural networks tend to use them as a black box tool. They do not access all of the possible variations, nor take into account the respective classification accuracies and costs. In our work, we focus on multiclass image classification, and in this research, we shed light on the trade-offs between systems using a single multiclass classification and multiple binary classifiers, respectively. We have tested these classifiers on several modern neural network architectures, including DenseNet, Inception v3, Inception ResNet v2, Xception, NASNet and MobileNet. We have compared several aspects of the performance of these architectures during training and testing using both classification styles in terms of classification speed and several classification accuracy metrics. Here, we present the results from experiments on a total of 99 networks: 11 multiclass and 88 individual binary networks, for an 8-class classification of medical images. In short, using multiple binary classification networks resulted in a more robust model (less variance) for the task at hand. However, on average, such a multi-network style performed the classification 7.6 times slower compared to a single network multiclass implementation. These collective findings show that both approaches can be applied to modern neural network structures. Several binary networks can often give more robust and increased classification accuracy, but at the cost of classification speed and resources consumption. Tor Jan Derek Berstad, Michael Riegler 0001, Håvard Espeland, Thomas de Lange, Pia H. Smedsrud, Konstantin Pogorelov, Håkon Kvale Stensland, Pål Halvorsen |
ISM | 7 |
| 2017 | A Holistic Multimedia System for Gastrointestinal Tract Disease DetectionabstractAnalysis of medical videos for detection of abnormalities and diseases requires both high precision and recall, but also real-time processing for live feedback and scalability for massive screening of entire populations. Existing work on this field does not provide the necessary combination of retrieval accuracy and performance.; [email protected] this paper, a multimedia system is presented where the aim is to tackle automatic analysis of videos from the human gastrointestinal (GI) tract. The system includes the whole pipeline from data collection, processing and analysis, to visualization. The system combines filters using machine learning, image recognition and extraction of global and local image features. Furthermore, it is built in a modular way so that it can easily be extended. At the same time, it is developed for efficient processing in order to provide real-time feedback to the doctors. Our experimental evaluation proves that our system has detection and localisation accuracy at least as good as existing systems for polyp detection, it is capable of detecting a wider range of diseases, it can analyze video in real-time, and it has a low resource consumption for scalability. Konstantin Pogorelov, Sigrun Losada Eskeland, Thomas de Lange, Carsten Griwodz, Kristin Ranheim Randel, Håkon Kvale Stensland, Duc-Tien Dang-Nguyen, Concetto Spampinato, Dag Johansen, Michael Riegler 0001, Pål Halvorsen |
MMSys | 6 |
| 2017 | Load Balancing of Multimedia Workloads for Energy Efficiency on the Tegra K1 Multicore ArchitectureabstractEnergy efficiency is a timely topic for modern mobile computing. Reducing the energy consumption of devices not only increases their battery lifetime, but also reduces the risk of hardware failure. Many researchers strive to understand the relationship between software activity and hardware power usage. A recurring strategy for saving power is to reduce operating frequencies. It is widely acknowledged that standard frequency scaling algorithms generally overreact to changes in hardware utilisation. More recent and original efforts attempt to balance software workloads on heterogeneous multicore architectures, such as the Tegra K1, which includes a quad-core CPU and a CUDA-capable GPU. However, it is not known whether it is possible to utilise these processor elements in parallel to save energy. Research into these types of systems are unfortunately often evaluated with the Performance Per Watt (PPW) metric, which is an unaccurate method because it ignores constant power usage from idle components. We show that this metric can end up increase energy usage on the Tegra K1, and give a false impression of how such systems consume energy. In reality, we show that it is much harder to save energy by balancing workloads between the heterogeneous cores of the Tegra K1, where we demonstrate only a 5% energy saving by offloading 10% DCT workload from the GPU to the CPU. Significantly more energy can be saved (up to 50 %) using the appropriate processor for different workloads. Kristoffer Robin Stokke, Håkon Kvale Stensland, Carsten Griwodz, Pål Halvorsen |
MMSys | 2 |
| 2016 | Efficient processing of videos in a multi-auditory environment using device lending of GPUsabstractIn this paper, we present a demo that utilizes Device Lending via PCI Express (PCIe) in the context of a multi-auditory environment. Device Lending is a transparent, low-latency cross-machine PCIe device sharing mechanism without any the need for implementing application-specific distribution mechanisms. As workload, we use a computer-aided diagnosis system that is used to automatically find polyps and mark them for medical doctors during a colonoscopy. We choose this scenario because one of the main requirements is to perform the analysis in real-time. The demonstration consists of a setup of two computers that demonstrates how Device Lending can be used to improve performance, as well as its effect of providing the performance needed for real-time feedback. We also present a performance evaluation that shows its real-time capabilities of it. Konstantin Pogorelov, Michael Riegler 0001, Jonas Markussen, Håkon Kvale Stensland, Pål Halvorsen, Carsten Griwodz, Sigrun Losada Eskeland, Thomas de Lange |
MMSys | 4 |
| 2016 | Right inflight?: a dataset for exploring the automatic prediction of movies suitable for a watching situationabstractIn this paper, we present the dataset Right Inflight developed to support the exploration of the match between video content and the situation in which that content is watched. Specifically, we look at videos that are suitable to be watched on an airplane, where the main assumption is that that viewers watch movies with the intent of relaxing themselves and letting time pass quickly, despite the inconvenience and discomfort of flight. The aim of the dataset is to support the development of recommender systems, as well as computer vision and multimedia retrieval algorithms capable of automatically predicting which videos are suitable for inflight consumption. Our ultimate goal is to promote a deeper understanding of how people experience video content, and of how technology can support people in finding or selecting video content that supports them in regulating their internal states in certain situations. Right Inflight consists of 318 human-annotated movies, for which we provide links to trailers, a set of pre-computed low-level visual, audio and text features as well as user ratings. The annotation was performed by crowdsourcing workers, who were asked to judge the appropriateness of movies for inflight consumption. Michael Riegler 0001, Martha A. Larson, Concetto Spampinato, Pål Halvorsen, Mathias Lux, Jonas Markussen, Konstantin Pogorelov, Carsten Griwodz, Håkon Kvale Stensland |
MMSys | 9 |
| 2016 | Computer aided disease detection system for gastrointestinal examinationsabstractIn this paper, we present the computer-aided diagnosis part of the EIR system [9], which can support medical experts in the task of detecting diseases and anatomical landmarks in the gastrointestinal (GI) system. This includes automatic detection of important findings in colonoscopy videos and marking them for the doctors. EIR is designed in a modular way so that it can easily be extended for other diseases. For this demonstration, we will focus on polyp detection, as our system is trained with the ASU-Mayo Clinic polyp database [5]. Michael Riegler 0001, Konstantin Pogorelov, Jonas Markussen, Mathias Lux, Håkon Kvale Stensland, Thomas de Lange, Carsten Griwodz, Pål Halvorsen, Dag Johansen, Peter Thelin Schmidt, Sigrun Losada Eskeland |
MMSys | 5 |
| 2016 | A high-precision, hybrid GPU, CPU and RAM power model for generic multimedia workloadsabstractEnergy efficiency of multimedia processing is a hot topic in modern, mobile computing where the lifetime of battery-powered devices is low. Authors often use power models as tools to evaluate the energy-efficiency of multimedia workloads and processing schemes. A challenge with these models is that they are built without sufficiently deep hardware knowledge and as a result they have the potential to mispredict substantially depending on hardware configuration. Typical rate-based power models can for example mispredict up to 70 % on the Tegra K1 SoC. Inspired by multimedia workloads, we introduce a modelling methodology which can be used to build a generic, high-precision power model for the Tegra K1's GPU and memory. By considering hardware utilisation, rail voltages, leakage currents and clocks, the model achieves an average accuracy above 99 % over all operating frequencies, and has been rigorously tested on several multimedia workloads. Our method exposes detailed insight into hardware and how it consumes energy. This knowledge is not only useful for researchers to understand how power models should be built, but also helps to understand what developers can do to minimise power usage. For example, experiments show that for a DCT benchmark, 3 % power can be saved by utilising non-coherent caches and smaller datatypes. Kristoffer Robin Stokke, Håkon Kvale Stensland, Carsten Griwodz, Pål Halvorsen |
MMSys | 2 |
| 2016 | Device lending in PCI express networksabstractThe challenge of scaling IO performance of multimedia systems to demands of their users has attracted much research. A lot of effort has gone into development of distributed systems that add little latency and computing overhead. For machines in PCI Express (PCIe) clusters, we propose Device Lending as a novel solution which works at a system level. Lars Bjørlykke Kristiansen, Jonas Markussen, Håkon Kvale Stensland, Michael Riegler 0001, Hugo Kohmann, Friedrich Seifert, Roy Nordstrøm, Carsten Griwodz, Pål Halvorsen |
NOSSDAV | 3 |
| 2015 | Scaling virtual camera services to a large number of usersabstractBy processing video footage from a camera array, one can easily make wide-field-of-view panorama videos. From the single panorama video, one can further generate multiple virtual cameras supporting personalized views to a large number of users based on only the few physical cameras in the array. However, giving personalized services to large numbers of users potentially introduces both bandwidth and processing bottlenecks, depending on where the virtual camera is processed. Vamsidhar Reddy, Ragnar Langseth, Håkon Kvale Stensland, Carsten Griwodz, Pål Halvorsen, Dag Johansen |
MMSys | 3 |
| 2015 | Energy efficient video encoding using the tegra K1 mobile processorabstractEnergy consumption is an important concern for mobile devices, where the evolution in battery storage capacity has not followed the power usage requirements of modern hardware. However, innovative and flexible hardware platforms give developers better means of optimising the energy consumption of their software. For example, the Tegra K1 System-on-Chip (SoC) offers two CPU clusters, GPU offloading, frequency scaling and other mechanisms to control the power and performance of applications. In this demonstration, the scenario is live video encoding, and participants can experiment with power usage and performance using the Tegra K1's hardware capabilities. A popular power-saving approach is a "race to sleep" strategy where the highest CPU frequency is used while the CPU has work to do, and then the CPU is put to sleep. Our own experiments indicate that an energy reduction of 28 % can be achieved by running the video encoder on the lowest CPU frequency at which the platform achieves an encoding frame rate equal to the minimum frame rate of 25 Frames Per Second (FPS). Kristoffer Robin Stokke, Håkon Kvale Stensland, Carsten Griwodz, Pål Halvorsen |
MMSys | 2 |
| 2014 | An Evaluation of Debayering Algorithms on GPU for Real-Time Panoramic Video RecordingabstractModern video cameras normally only capture a single color per pixel, commonly arranged in a Bayer pattern. This means that we must restore the missing color channels in the image or the video frame in post-processing, a process referred to as debayering. In a live video scenario, this operation must be performed efficiently in order to output each frame in real-time, while also yielding acceptable visual quality. Here, we evaluate debayering algorithms implemented on a GPU for real-time panoramic video recordings using multiple 2K-resolution cameras. Ragnar Langseth, Vamsidhar Reddy, Håkon Kvale Stensland, Carsten Griwodz, Pål Halvorsen |
ISM | 3 |
| 2014 | Using a Commodity Hardware Video Encoder for Interactive Video StreamingabstractOver the last years, video streaming has become one of the most dominant Internet services. A trend now is that due to the increased availability of high-speed internet access, multimedia services are becoming more interactive and immersive. Examples of such applications are both cloud gaming and systems where users can interact with high-resolution content. Over the last few years, hardware video encoders have been built into commodity hardware. We evaluate one of these encoders in a scenario where we have individual streams delivered to the end users. Our results show that we can reduce almost half of the CPU time spent on video processing, while also greatly reducing the power consumption on the system. We also compare the visual video quality and the frame size of the hardware based encoder, and we find no significant difference compared to a software based approach. Martin Alexander Wilhelmsen, Håkon Kvale Stensland, Vamsidhar Reddy, Asgeir Mortensen, Ragnar Langseth, Carsten Griwodz, Pål Halvorsen |
ISM | 2 |
| 2014 | Automatic Real-Time Zooming and Panning on Salient Objects from a Panoramic VideoabstractThe proposed demo shows how our system automatically zooms and pans into tracked objects in panorama videos. At the conference site, we will set up a two-camera version of the system, generating live panorama videos, where the system zooms and pans tracking people using colored hats. Additionally, using a stored soccer game video from a five 2K camera setup at Alfheim stadium in Tromsø from the European league game between Tromsø IL and Tottenham Hotspurs, the system automatically follows the ball. Vamsidhar Reddy, Ragnar Langseth, Håkon Kvale Stensland, Carsten Griwodz, Pål Halvorsen, Øystein Landsverk |
ACM Multimedia | 3 |
| 2014 | Real-Time HDR Panorama VideoabstractThe interest for wide field of view panorama video is increasing. In this respect, we have an application that uses an array of cameras that overlook a soccer stadium. The input of these cameras are stitched together to provide a panoramic view of the stadium. One of the challenges we face is that large parts of the field are obscured by shadows on sunny days. Such circumstances cause unsatisfying video quality. We have therefore implemented and evaluated multiple algorithms related to high dynamic range (HDR) video. The evaluation shows that a combination of several approaches gives the most useful results in our scenario. Lorenz Kellerer, Vamsidhar Reddy, Ragnar Langseth, Håkon Kvale Stensland, Carsten Griwodz, Dag Johansen, Pål Halvorsen |
ACM Multimedia | 4 |
| 2014 | Be your own cameraman: real-time support for zooming and panning into stored and live panoramic videoabstractHigh-resolution panoramic video with a wide field-of-view is popular in many contexts. However, in many examples, like surveillance and sports, it is often desirable to zoom and pan into the generated video. A challenge in this respect is real-time support, but in this demo, we present an end-to-end real-time panorama system with interactive zoom and panning. Our system installed at Alfheim stadium, a Norwegian premier league soccer team, generates a cylindrical panorama from five 2K cameras live where the perspective is corrected in real-time when presented to the client. This gives a better and more natural zoom compared to existing systems using perspective panoramas and zoom operations using plain crop. Our experimental results indicate that virtual views can be generated far below the frame-rate threshold, i.e., on a GPU, the processing requirement per frame is about 10 milliseconds. Vamsidhar Reddy, Ragnar Langseth, Håkon Kvale Stensland, Pierre Gurdjos, Vincent Charvillat, Carsten Griwodz, Dag Johansen, Pål Halvorsen |
MMSys | 3 |
| 2014 | Automatic event extraction and video summaries from soccer gamesabstractBagadus is a prototype of a soccer analysis application which integrates a sensor system, a video camera array and soccer analytics annotations. The current prototype is installed at Alfheim Stadium in Norway, and provides a large set of new functions compared to existing solutions. One important feature is to automatically extract video events and summaries from the games, i.e., an operation that traditionally consumes a huge amount of time. In this demo, we demonstrate how our integration of subsystems enable several types of summaries to be generated automatically, and we show that the video summaries are displayed with a response time around one second. Asgeir Mortensen, Vamsidhar Reddy, Håkon Kvale Stensland, Carsten Griwodz, Dag Johansen, Pål Halvorsen |
MMSys | 3 |
| 2014 | Soccer video and player position datasetabstractThis paper presents a dataset of body-sensor traces and corresponding videos from several professional soccer games captured in late 2013 at the Alfheim Stadium in Tromsø, Norway. Player data, including field position, heading, and speed are sampled at 20Hz using the highly accurate ZXY Sport Tracking system. Additional per-player statistics, like total distance covered and distance covered in different speed classes, are also included with a 1Hz sampling rate. The provided videos are in high-definition and captured using two stationary camera arrays positioned at an elevated position above the tribune area close to the center of the field. The camera array is configured to cover the entire soccer field, and each camera can be used individually or as a stitched panorama video. This combination of body-sensor data and videos enables computer-vision algorithms for feature extraction, object tracking, background subtraction, and similar, to be tested against the ground truth contained in the sensor traces. Svein Arne Pettersen, Dag Johansen, Håvard D. Johansen, Vegard Berg-Johansen, Vamsidhar Reddy, Asgeir Mortensen, Ragnar Langseth, Carsten Griwodz, Håkon Kvale Stensland, Pål Halvorsen |
MMSys | 9 |
| 2014 | Bagadus: An integrated real-time system for soccer analyticsabstractThe importance of winning has increased the role of performance analysis in the sports industry, and this underscores how statistics and technology keep changing the way sports are played. Thus, this is a growing area of interest, both from a computer system view in managing the technical challenges and from a sport performance view in aiding the development of athletes. In this respect, Bagadus is a real-time prototype of a sports analytics application using soccer as a case study. Bagadus integrates a sensor system, a soccer analytics annotations system, and a video processing system using a video camera array. A prototype is currently installed at Alfheim Stadium in Norway, and in this article, we describe how the system can be used in real-time to playback events. The system supports both stitched panorama video and camera switching modes and creates video summaries based on queries to the sensor system. Moreover, we evaluate the system from a systems point of view, benchmarking different approaches, algorithms, and trade-offs, and show how the system runs in real time. Håkon Kvale Stensland, Vamsidhar Reddy, Marius Tennøe, Espen Helgedagsrud, Mikkel Næss, Henrik Kjus Alstad, Asgeir Mortensen, Ragnar Langseth, Sigurd Ljødal, Øystein Landsverk, Carsten Griwodz, Pål Halvorsen, Magnus Stenhaug, Dag Johansen |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2013 | Efficient Implementation and Processing of a Real-Time Panorama Video PipelineabstractHigh resolution, wide field of view video generated from multiple camera feeds has many use cases. However, processing the different steps of a panorama video pipeline in real-time is challenging due to the high data rates and the stringent requirements of timeliness. We use panorama video in a sport analysis system where video events must be generated in real-time. In this respect, we present a system for real-time panorama video generation from an array of low-cost CCD HD video cameras. We describe how we have implemented different components and evaluated alternatives. We also present performance results with and without co-processors like graphics processing units (GPUs), and we evaluate each individual component and show how the entire pipeline is able to run in real-time on commodity hardware. Marius Tennøe, Espen Helgedagsrud, Mikkel Næss, Henrik Kjus Alstad, Håkon Kvale Stensland, Vamsidhar Reddy, Dag Johansen, Carsten Griwodz, Pål Halvorsen |
ISM | 5 |
| 2013 | Demonstrating Hundreds of AIs in One Scene
Kjetil Raaen, Andreas Petlund, Håkon Kvale Stensland |
ICEC | 3 |
| 2013 | Bagadus: an integrated system for arena sports analytics: a soccer case studyabstractSports analytics is a growing area of interest, both from a computer system view to manage the technical challenges and from a sport performance view to aid the development of athletes. In this paper, we present Bagadus, a prototype of a sports analytics application using soccer as a case study. Bagadus integrates a sensor system, a soccer analytics annotations system and a video processing system using a video camera array. A prototype is currently installed at Alfheim Stadium in Norway, and in this paper, we describe how the system can follow and zoom in on particular player(s). Next, the system will playout events from the games using stitched panorama video or camera switching mode and create video summaries based on queries to the sensor system. Furthermore, we evaluate the system from a systems point of view, benchmarking different approaches, algorithms and tradeoffs. Pål Halvorsen, Simen Saegrov, Asgeir Mortensen, David K. C. Kristensen, Alexander Eichhorn, Magnus Stenhaug, Stian Dahl, Håkon Kvale Stensland, Vamsidhar Reddy, Carsten Griwodz, Dag Johansen |
MMSys | 8 |
| 2011 | Improved Multi-Rate Video EncodingabstractAdaptive HTTP streaming is frequently used for both live and on-Demand video delivery over the Internet. Adaptive ness is often achieved by encoding the video stream in multiple qualities (and thus bit rates), and then transparently switching between the qualities according to the bandwidth fluctuations and the amount of resources available for decoding the video content on the end device. For this kind of video delivery over the Internet, H.264 is currently the most used codec, but VP8 is an emerging open-source codec expected to compete with H.264 in the streaming scenario. The challenge is that, when encoding video for adaptive video streaming, both VP8 and H.264 run once for each quality layer, i.e., consuming both time and resources, especially important in a live video delivery scenario. In this paper, we address the resource consumption issues by proposing a method for reusing redundant steps in a video encoder, emitting multiple outputs with varying bit rates and qualities. It shares and reuses the computational heavy analysis step, notably macro-block mode decision, intra prediction and inter prediction between the instances, and outputs video in several rates. The method has been implemented in the VP8 reference encoder, and experimental results show that we can encode the different quality layers at the same rates and qualities compared to the VP8 reference encoder, while reducing the encoding time significantly. Dag Haavi Finstad, Håkon Kvale Stensland, Håvard Espeland, Pål Halvorsen |
ISM | 2 |
| 2011 | Processing of multimedia data using the P2G frameworkabstractIn this demo, we present the P2G framework designed for processing distributed real-time multimedia data. P2G supports arbitrarily complex dependency graphs with cycles, branches and deadlines. P2G is implemented to scale transparently with available resources, i.e., a concept familiar from the cloud computing paradigm. Additionally, P2G supports heterogeneous computing resources, such as x86 and GPU processing cores. We have implemented an interchangeable P2G kernel language which is meant to expose fundamental concepts of the P2G programming model and ease the application development. Here, we demonstrate the P2G execution node using a MJPEG encoder as an example workload when dynamically adding and removing processing cores. Paul B. Beskow, Håkon Kvale Stensland, Håvard Espeland, Espen A. Kristiansen, Preben N. Olsen, Ståle Kristoffersen, Carsten Griwodz, Pål Halvorsen |
ACM Multimedia | 2 |
| 2010 | Tips, tricks and troubles: optimizing for cell and GPUabstractWhen used efficiently, modern multicore architectures, such as Cell and GPUs, provide the processing power required by resource demanding multimedia workloads. However, the diversity of resources exposed to the programmers, intrinsically requires specific mindsets for efficiently utilizing these resources - not only compared to an x86 architecture, but also between the Cell and the GPUs. In this context, our analysis of 14 different Motion-JPEG implementations indicates that there exists a large potential for optimizing performance, but there are also many pitfalls to avoid. By experimentally evaluating algorithmic choices, inter-core data communication (memory transfers) and architecture-specific capabilities, such as instruction sets, we present tips, tricks and troubles with respect to efficient utilization of the available resources. Håkon Kvale Stensland, Håvard Espeland, Carsten Griwodz, Pål Halvorsen |
NOSSDAV | 1 |
| 2009 | Improving file tree traversal performance by scheduling I/O operations in user spaceabstractCurrent in-kernel disk schedulers provide efficient means to optimize the order (and minimize disk seeks) of issued, in-queue I/O requests. However, they fail to optimize sequential multi-file operations, like traversing a large file tree, because only requests from one file are available in the scheduling queue at a time. We have therefore investigated a user-level, I/O request sorting approach to reduce inter-file disk arm movements. This is achieved by allowing applications to utilize the placement of inodes and disk blocks to make a one sweep schedule for all file I/Os requested by a process, i.e., data placement information is read first before issuing the low-level I/O requests to the storage system. Our experiments with a modified version of tar show reduced disk arm movements and large performance improvements. Carl Henrik Lunde, Håvard Espeland, Håkon Kvale Stensland, Pål Halvorsen |
IPCCC | 3 |
| 2008 | Making an SCI fabric dynamically fault tolerantabstractIn this paper we present a method for dynamic fault tolerant routing for SCI networks implemented on Dolphin Interconnect Solutions hardware. By dynamic fault tolerance, we mean that the interconnection network reroutes affected packets around a fault, while the rest of the network is fully functional. To the best of our knowledge this is the first reported case of dynamic fault tolerant routing available on commercial off the shelf interconnection network technology without duplicating hardware resources. The development is focused around a 2-D torus topology, and is compatible with the existing hardware, and software stack. We look into the existing mechanisms for routing in SCI. We describe how to make the nodes that detect the faulty component do routing decisions, and what changes are needed in the existing routing to enable support for local rerouting. The new routing algorithm is tested on clusters with real hardware. Our tests show that distributed databases like MySQL can run uninterruptedly while the network reacts to faults. The solution is now part of Dolphin Interconnect Solutions SCI driver, and hardware development to further decrease the reaction time is underway. Håkon Kvale Stensland, Olav Lysne, Roy Nordstrøm, Hugo Kohmann |
IPDPS | 1 |
| 2008 | Transparent protocol translation and load balancing on a network processor in a media streaming scenarioabstractToday, major newspapers and TV stations make live and on-demand audio/video content available, video-on-demand services are becoming common and even personal media are frequently uploaded to streaming sites. The discussion about the best transport protocol for streaming has been going on for years. Currently, HTTP-streaming is usual although the transport of streaming media data over TCP is hindered by TCP's probing behavior, which results in the rapid reduction and slow recovery of the packet rates. On the other hand, UDP has been criticized for being unfair against TCP, and it is therefore often blocked by access network providers. Håvard Espeland, Carl Henrik Lunde, Håkon Kvale Stensland, Carsten Griwodz, Pål Halvorsen |
NOSSDAV | 3 |
| 2008 | Evaluation of multi-core scheduling mechanisms for heterogeneous processing architecturesabstractGeneral-purpose CPUs with multiple cores are established products, and new heterogeneous technology like the Cell broadband engine and general-purpose GPUs bring an even higher degree of true multi-processing into the market. However, means for utilizing the processing power is immature. Current tools typically assume that exclusive use of these resources is sufficient, but this assumption will soon be invalid because the interest in using their processing power for general-purpose tasks. Among the applications that can benefit from such technology is transcoding support for distributed media applications, where remote participants join and leave dynamically. Transcoding consists of several clearly separated processing operations that consume a lot of resources, such that individual processing units are unable to handle all operations of a session of arbitrary size. The individual operations can then be distributed over several processing units, and data must be moved between them according to the dependencies between operations. Many multi-processor scheduling approaches exist, but to the best of our knowledge, a challenge is still to find mechanisms that can schedule dynamic workloads of communicating operations while taking both the processing and communication requirements into account. For such applications, we believe that feasible scheduling can be performed in two levels, i.e., divided into the task of placing a job onto a processing unit and the task of multitasking time-slices within a single processing unit. We have implemented some simple high-level scheduling mechanisms and simulated a video conferencing scenario running on topologies inspired by existing systems from Intel, AMD, IBM and nVidia. Our results show the importance of using an efficient high-level scheduler. Håkon Kvale Stensland, Carsten Griwodz, Pål Halvorsen |
NOSSDAV | 1 |
| 2007 | Transparent protocol translation for streamingabstractThe transport of streaming media data over TCP is hindered by TCP's probing behavior that results in the rapid reduction and slow recovery of the packet rates. On the other side, UDP has been criticized for being unfair against TCP connections, and it is therefore often blocked out in the access networks. In this paper, we try to benefit from a combined approach using a proxy that transparently performs transport protocol translation. We translate HTTP requests by the client transparently into RTSP requests, and translate the corresponding RTP/UDP/AVP stream into the corresponding HTTP response. This enables the server to use UDP on the server side and TCP on the client side. This is beneficial for the server side that scales to a higher load when it doesn't have to deal with TCP. On the client side, streaming over TCP has the advantage that connections can be established from the client side, and data streams are passed through firewalls. Preliminary tests demonstrate that our protocol translation delivers a smoother stream compared to HTTP-streaming where the TCP bandwidth oscillates heavily. Håvard Espeland, Carl Henrik Lunde, Håkon Kvale Stensland, Carsten Griwodz, Pål Halvorsen |
ACM Multimedia | 3 |