Otto J. Anshus

dblp:95/6470 · also Otto Johan Anshus · DBLP profile ↗
← Back
33ranked-venue papers
0as first author
8since 2021 · last 2025
0000-0003-3579-6566ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 21 · 6 since 2021Software engineering, systems software and programming languages · 3Human-computer interaction and ubiquitous computing · 3Computer networks · 2 · 1 since 2021Artificial intelligence and machine learning · 1Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2025 Using Analytics to Predict Data Dissemination Policy Under Energy and Coverage Constraints
abstract
Disseminating data in a wireless distributed system under limited resources, where nodes have constrained energy and communications capabilities, is challenging. In such contexts, common data dissemination techniques cannot be used. Instead, simple device-to-device communication policies allow mitigating the impact of communications on nodes energy consumption. However, depending on nodes configuration (up-times duration, wireless technology capabilities and energy consumption), choosing a suitable communication policy is challenging.In this paper, we propose and study two approaches based on classification algorithms that aim at predicting the most suitable communication policy to use, for a given node configuration, to match a given coverage and energy consumption target. The first approach called in situ learning, trains the classification models during deployment. The second approach called offline learning, uses existing data that are collected from previous deployments or simulations. Results show that, for a resource constrained environment, common classification models can take several months to converge with in situ learning. Depending on the policy, in situ learning can have a significant energy consumption overhead. Results underline offline learning as an interesting alternative to in situ learning, but requires to collect data from previous deployments or simulations.
Loïc Guegan, Issam Raïs, Otto J. Anshus
MSWiM3
2025 A large-scale study of the impact of node behavior on loosely coupled data dissemination: The case of the distributed Arctic observatory
abstract
A Cyber-Physical System (CPS) deployed in remote and resource-constrained environments faces multiple challenges. It has, no or limited: network coverage, possibility of energy replenishment, physical access by humans. Cyber-physical nodes deployed to observe and interact with the Arctic tundra face these challenges. They are subject to environmental factors such as avalanches, low temperatures, snow, ice, water and wild animals. Without energy supply infrastructures and humans available, nodes must achieve long operational lifetime from a single battery charge. They must be extremely energy-efficient. To reduce energy costs and increase their energy efficiency, cyber-physical nodes sleep most of the time, and avoid to communicate when they are unreachable. But, a CPS needs to disseminate data between the nodes for multiple purposes including data reporting to a back-end service, resilient operations, safe-keeping of observational data, and propagating nodes updates. Loosely-coupled data dissemination policies offer this possibility [1] . Although, investigations should be made on their applicability to large-scale CPS. In this paper, we evaluate and discuss the efficiency in energy, time and number of successful delivery of four data dissemination policies proposed in [1] . This evaluation is based on flow-level simulations. We study small and large-scale CPS, and evaluate the effects of the number of nodes and the size of the disseminated data on the nodes energy consumption and the dissemination's delivery success. To mitigate negative effects raised on large-scale CPS and large disseminated data sizes, different strategies are proposed and evaluated. We show that energy saving strategies do not always imply energy efficiency, and better data dissemination often comes at a cost. This last result highlights the importance of simulation prior to real CPS deployments in constrained environments. • CPS deployed in remote and resource-constrained environments faces challenges. • Energy efficient and scalable approaches for data dissemination in CPS are needed. • Simulation of CPS provides trade-offs in energy and data dissemination performance.
Loïc Guegan, Issam Raïs, Otto J. Anshus
J. Parallel Distributed Comput.3
2022 Impact of loosely coupled data dissemination policies for resource challenged environments
abstract
A Cyber-Physical System (CPS) deployed to a resource-constrained environment can face multiple challenges like (i) no or limited network coverage, (ii) no or limited possibility of energy replenishment, (iii) no or limited physical access by humans, (iv) nodes must cope with environmental factors including avalanches, low temperatures, snow, ice, water and wild animals. Devices being part of such a CPS must be battery powered and be energy efficient to achieve long life-time. However, the CPS still has to disseminate data to increase resiliency, safely keep results or update nodes. A trade-off between energy spent and data dissemination needs to be found, ideally by minimizing the energy usage and maximizing the relevant data dissemination performance metrics. In this paper, we evaluate and discuss the efficiency (in energy, time and number of successful distributions) of multiple data distribution policies by mean of flow-level simulations. We report on the trade-off between (i) successful data dissemination and (ii) energy and uptime overheads resulting from the usage of loosely coupled policies. To fully explore the scope of possibilities, we simulate a wide range of scenarios extracted from real measurements and previous deployments. Characteristics of CPS devices developed by the Distributed Arctic Observatory (DAO) are used as simulation platforms. Results show that an efficient policy in a given scenario can perform worse in another scenario. We also show that simple policies, especially when combined, can help in minimizing the energy consumed by most of the devices composing the CPS and maximizing the relevant dissemination performance metrics.
Issam Raïs, Loïc Guegan, Otto J. Anshus
CCGRID3
2022 Detecting the Arrival of an Update at Mostly Sleeping Cyber-Physical IoT Nodes
abstract
IoT and cyber-physical nodes with long operational lifetimes will be updated with some frequency after deployment. However, nodes deployed to environments like the arctic tundra are faced with several resource constraints, including a very small energy budget. To reduce energy consumption, the nodes are mostly sleeping. This makes updates more difficult because the update procedure can be interrupted by the nodes deciding to sleep or shut down. In addition, the updating system itself should be energy efficient.We report on an approach and a prototype system for a single phase of the updating, concerned only with the problem of detecting that an update has arrived locally at a node.The nodes can expect to receive updates at different update schedules. A series of performance measuring experiments were conducted on the update detection system to document how it behaves for different update schedules.We conclude that to find a sweet spot for different update schedules with regard to the number of activations of the detection mechanism, and by implication also the energy consumption, and the time it takes before and update is detected, the update detection system must adapt to the different update schedules. While several factors impact this, the most significant effect comes from varying the time between activations of the detection mechanism.
Roberth Tollefsen, Otto J. Anshus
DCOSS2
2022 LoRaLitE: LoRa protocol for Energy-Limited environments
abstract
To do in-situ observations of the arctic tundra, many small sensor nodes are used. Reporting the data to remote back-ends is hard because backhaul networks are scarce, and a node cannot expect to see one. Also, there are typically no humans physically close to the nodes to fetch the data and to replace batteries. Consequently, the nodes need a highly available, energy-efficient long range network. While LoRaWAN provides energy-efficient communication for nodes, it does so at the cost of needing an always-on gateway with a high energy consumption. The gateway is also a single point of failure. Instead we propose LoRaLitE for Energy-Limited environments where the gateway enters sleep phases to reduce energy consumption. The sleep phases of the nodes are coordinated accordingly. For high availability, any end-node with a backhaul-network can be elected to become the gateway. We conducted a series of simulation experiments to document the performance behavior of both LoRaWanand LoRaLitE. The results show that the LoRaLitE gateway spends from 10 to 10,000 times less energy than a LoRaWAN gateway. A LoRaLitE node spends from 13% to 42% more energy than a LoRaWAN node. The achievable bandwidth for LoRaLitE is only insignificantly smaller than for LoRaWAN. A LoRaWangateway needs a relatively large, heavy battery. If the gateway fails, assuming a new gateway can be elected, several nodes must have similar large batteries to be able to take over as the gateway. This is impractical to achieve for more than a few nodes because the nodes become more expensive, harder to camouflage, and impractical to deploy. A LoRaLitE gateway on the other hand only needs a battery like any other node. Even if all nodes need larger batteries than for the LoRaWancase, the increase is in practical terms insignificant. We conclude that LoRaLitE makes it realistic to use LoRa for nodes on the arctic tundra, providing for insignificant increased energy usage at the nodes, insignificant reduced bandwidth, but with significant increase in gateway availability. LoRaLitE also provides for an easier deployment in practice because all nodes can be kept small and light,
Lukasz Sergiusz Michalik, Loïc Guegan, Issam Raïs, Otto J. Anshus, John Markus Bjørndalen
MASCOTS4
2021 Experiences Building and Deploying Wireless Sensor Nodes for the Arctic Tundra
abstract
The arctic tundra is most sensitive to climate change. The change can be quantified from observations of the fauna, flora and weather conditions. To do observations at sufficient spatial and temporal resolution, ground-based observation nodes with sensors are needed.However, the arctic tundra is resource-limited with regards to energy, data networks, and humans. There are also regulatory and practical obstacles. Consequently, observation nodes must be small and unobtrusive, have a year or longer operational lifetime from small batteries, and be able to report results and receive software updates over scarce back-haul networks.We describe the architecture, design, and implementation of prototype observation nodes deployed to the arctic tundra for the periods August 2019 to July 2020 and August 2020 to July 2021.For the 2019 deployment, ten nodes were each placed inside ten existing camera traps. A camera trap is a box with a wildlife camera taking pictures of rodents when they enter the box from tunnels under snow and ice. For the 2020 deployment, eight nodes were located pairwise inside four camera traps.Each node measures carbon dioxide level and temperature inside the camera trap during the winter season. A node reports its state and observational data each night over a commercial low power IoT telecom back-haul network, if available.We report on the issues encountered doing actual deployments of the prototype nodes. For each issue, we describe the reason for why it happened, relate it to the architecture, design and implementation, and explain what we did about it.
Michael J. Murphy, Øystein Tveito, Eivind Flittie Kleiven, Issam Raïs, Eeva M. Soininen, John Markus Bjørndalen, Otto J. Anshus
CCGRID7
2021 Distribution of Updates to IoT Nodes in a Resource-Challenged Environment
abstract
IoT nodes need to be updated after deployment. However, doing so for nodes deployed to resource-challenged environments, like the arctic tundra, is a challenge. Because humans as the common case cannot physically visit the nodes, updating must be done from a remote update service over a back-haul data network. However, most nodes are not in range of a back-haul data network. Even when nodes are in range, they probably sleep to conserve energy and the remote cloud back-end therefore cannot communicate with them.We report on an approach and a prototype system for distributing updates from a cloud update distribution service to nodes. We assume that nodes carry one or several local area network technologies supporting short-lived point-to-point ad-hoc communication between two nodes at a time in range of each other. We assume that a single node in each neighborhood has a back-haul network, and delegate to this node to further distribute updates inside the neighborhood.A series of performance measuring experiments were conducted on the update distribution system when it executes on nodes with behaviors from always-on to mostly-off. We document how the distribution system behaves through a set of performance metrics. The results are very sensitive to the behavior of the nodes.
Roberth Tollefsen, Issam Raïs, John Markus Bjørndalen, Phuong Hoai Ha, Otto J. Anshus
CCGRID5
2021 Implementation and Performance Analysis for Energy Harvesting-based Wireless Sensor Networks
abstract
A wireless sensor network (WSN) depends on energy to operate and considerable efforts have been devoted to extend its operational lifetime. Energy harvesting (EH) has emerged as one promising approach to prolong the lifetime of a WSN without interrupting its operations. In an EH-WSN, sensor nodes typically have batteries being charged with energy harvested from ambient environments such as wind, solar, water mills etc. However, uncertainties remain about when and how much energy can be harvested over time from the environments and about traffic fluctuation at each node. Unfortunately, existing works on EH-WSN are limited in evaluating the impacts of these factors and the number of the nodes in an EH-WSN on its key performance metrics. This paper, therefore, presents a general method and a practical implementation of a node capable of harvesting solar energy in an EH-WSN. The paper then evaluates the performance of the EH-WSN. Explicit mathematical expressions define models for network operational time, throughput and reliability of the EH-WSN and simulations are used to validate the models. Influential factors are energy arrival rate, traffic rate at each node and the number of nodes in an EH-WSN. The results document that the analytical models and simulations correspond.
Nga Dinh, Øystein Tveito, Phuong Hoai Ha, Otto J. Anshus
IECON4
2020 Trading Data Size and CNN Confidence Score for Energy Efficient CPS Node Communications
abstract
In a context of Cyber-Physical Systems (CPS), energy-efficiency is a critical factor to achieve long operational life-time. The constraint of using battery-powered devices adds degrees of complexity, especially in a hard to reach environment with scarce network and energy resources. The reporting of data consumes large amount of energy, reducing the life-time of both individual nodes and the CPS as a whole. One way to reduce the energy cost of communication is to reduce the number of Bits to transmit. However, this is a viable approach only if the transmitted data remain suitable for further analysis.In this paper, we report on the effect of reducing the image dimensions on the confidence score computed by a convolutional neural network (CNN) determining the species of animals present in images. We also report on the energy consumption of transmitting full vs. reduced dimensions of images. CPS devices and CNNs developed by the Distributed Arctic Observatory (DAO) project are used as experimental platforms.The results show that the energy needed to report the images can be reduced by up to 98% while only reducing the average confidence of determining the species correctly by 0.10%.
Issam Raïs, Otto J. Anshus, John Markus Bjørndalen, Daniel Balouek-Thomert, Manish Parashar
CCGRID2
2019 UAVs as a Leverage to Provide Energy and Network for Cyber-Physical Observation Units on the Arctic Tundra
abstract
Observing an environment is essential to its understanding. This is certainly true for the arctic tundra as it is extremely sensitive to climate change. Satellites are used to observe large areas of the arctic tundra. However, the measurements are not at sufficient spatial resolution to be used to determine the species of animals, and to measure humidity, temperature, and CO2 levels for small areas. Ground-based measurements, which represent less that one percent of the experiments carried on the arctic tundra, must be done. Deploying and maintaining ground-based instruments is hard to do because visiting the arctic tundra is expensive, time consuming, and dangerous. Instead, automation is needed, where instruments can operate independently for long periods of time, and allow interaction with remote users for reporting of data and control of the instruments. However, the arctic tundra also has, as a common case, limited data network and energy coverage. We present the Distributed Arctic Observatory, where instruments called Observation Units (OUs) are distributed across the arctic tundra to measure a range of environmental state variables, and report them to where they are needed for further analysis. The DAO project uses Unmanned Aerial Vehicles (UAVs) to provide backhaul network access and energy to OUs. The modifications applied to the UAV and to OUs are presented. We describe the experiences gathered from practical use of the UAV to provide network access and energy to OUs located on the ground at 70°N. We show that it is feasible to apply UAVs to provide network access and energy to OUs on the arctic tundra, albeit practical restrictions.
Issam Raïs, John Markus Bjørndalen, Phuong Hoai Ha, Ken-Arne Jensen, Lukasz Sergiusz Michalik, Håvard Mjøen, Øystein Tveito, Otto J. Anshus
DCOSS8
2019 Efficient Concurrent Search Trees Using Portable Fine-Grained Locality
abstract
Concurrent search trees are crucial data abstractions widely used in many important systems such as databases, file systems and data storage. Like other fundamental abstractions for energy-efficient computing, concurrent search trees should support both high concurrency and fine-grained data locality in a platform-independent manner. However, existing portable fine-grained locality-aware search trees such as ones based on the van Emde Boas layout (vEB-based trees) poorly support concurrent update operations while existing highly-concurrent search trees such as non-blocking search trees do not consider fine-grained data locality. In this paper, we first present a novel methodology to achieve both portable fine-grained data locality and high concurrency for search trees. Based on the methodology, we devise a novel locality-aware concurrent search tree called GreenBST. To the best of our knowledge, GreenBST is the first practical search tree that achieves both portable fine-grained data locality and high concurrency. We analyze and compare GreenBST energy efficiency (in operations/Joule) and performance (in operations/second) with seven prominent concurrent search trees on a high performance computing (HPC) platform (Intel Xeon), an embedded platform (ARM), and an accelerator platform (Intel Xeon Phi) using parallel micro- benchmarks (Synchrobench). Our experimental results show that GreenBST achieves the best energy efficiency and performance on all the different platforms. GreenBST achieves up to 50 percent more energy efficiency and 60 percent higher throughput than the best competitor in the parallel benchmarks. These results confirm the viability of our new methodology to achieve both portable fine-grained data locality and high concurrency for search trees.
Phuong Hoai Ha, Otto J. Anshus, Ibrahim Umar
IEEE Trans. Parallel Distributed Syst.2
2017 Wait-Free Programming for General Purpose Computations on Graphics Processors
abstract
The fact that graphics processors (GPUs) are today’s most powerful computational hardware for the dollar has motivated researchers to utilize the ubiquitous and powerful GPUs for general-purpose computing. However, unlike CPUs, GPUs are optimized for processing 3D graphics (e.g., graphics rendering), a kind of data-parallel applications, and consequently, several GPUs do not support strong synchronization primitives to coordinate their cores. This prevents the GPUs from being deployed more widely for general-purpose computing. This paper aims at bridging the gap between the lack of strong synchronization primitives in the GPUs and the need for strong synchronization mechanisms in parallel applications. Based on the intrinsic features of typical GPU architectures, we construct strong synchronization objects such as wait-free and$t$-resilientread-modify-writeobjects for a general model of GPU architectures without hardware synchronization primitives such astest-and-setandcompare-and-swap. Accesses to the wait-free objects have time complexity$O(N)$, where$N$is the number of processes. The wait-free objects have the optimal space complexity$O(N^2)$. Our result demonstrates that it is possible to construct wait-free synchronization mechanisms for GPUs without strong synchronization primitives in hardware and that wait-free programming is possible for such GPUs.
Phuong Hoai Ha, Philippas Tsigas, Otto J. Anshus
IEEE Trans. Computers3
2016 GreenBST: Energy-Efficient Concurrent Search Tree
Ibrahim Umar, Otto J. Anshus, Phuong Hoai Ha
Euro-Par2
2016 Effect of portable fine-grained locality on energy efficiency and performance in concurrent search trees
abstract
Recent research has suggested that improving fine-grained data-locality is one of the main approaches to improving energy efficiency and performance. However, no previous research has investigated the effect of the approach on these metrices in the case of concurrent data structures.
Ibrahim Umar, Otto J. Anshus, Phuong Hoai Ha
PPoPP2
2015 DeltaTree: A Locality-aware Concurrent Search Tree
abstract
Like other fundamental abstractions for high-performance computing, search trees need to support both high concurrency and data locality. However, existing locality-aware search trees based on the van Emde Boas layout (vEB-based trees), poorly support concurrent (update) operations.
Ibrahim Umar, Otto J. Anshus, Phuong Hoai Ha
SIGMETRICS2
2014 Masking the Effects of Delays in Human-to-Human Remote Interaction
abstract
Humans can interact remotely with each other through computers.Systems supporting this include teleconferencing, games and virtual environments.There are delays from when a human does an action until it is reflected remotely.When delays are too large, they will result in inconsistencies in what the state of the interaction is as seen by each participant.The delays can be reduced, but they cannot be removed.When delays become too large the effects they create on the humanto-human remote interaction can be partially masked to achieve an illusion of insignificant delays.The MultiStage system is a human-to-human interaction system meant to be used by actors at remote stages creating a common virtual stage.Each actor is remotely represented by a remote presence created based on a stream of data continuously recorded about the actor and being sent to all stages.We in particular report on the subsystem of MultiStage masking the effects of delays.The most advanced masking approach is done by having each stage continuously look for late data, and when masking is determined to be needed, the system switches from using a live stream to a pre-recorded video of an actor.The system can also use a computable model of an actor creating a remote presence substituting for the live stream.The present prototype uses a simple human skeleton model.
John Markus Bjørndalen, Phuong Hoai Ha, Otto J. Anshus
FedCSIS4
2013 Accurate weather forecasting through locality based collaborative computing
abstract
The Collaborative Symbiotic Weather Forecasting (CSWF) system lets a user compute a short time, high-resolution forecast for a small region around the user, in a few minutes, on-demand, on a PC. A collaborated forecast giving better uncertainty estimation is then created using forecasts from other u
Bård Fjukstad, John Markus Bjørndalen, Otto J. Anshus
CollaborateCom3
2013 Global interaction space for user interaction with a room of computers
abstract
To interact with a computer, a user can walk up to it and interact with it through its local interaction space defined by its input devices. With multiple computers in a room, the user can walk up to each computer and interact with it. However, this can be logistically impractical and forces the user to learn each computers local interaction space. Interaction involving multiple computers also becomes hard or even impossible to do. We propose to have a global interaction space, letting users, through in-room gestures, select and issue commands to one or multiple computers in the room. A global interaction space has functionality to sense and record the state of a room, including location of computers, users, and gestures, and use this to issue commands to each computer. A prototype has been implemented in a room with multiple computers and a wall-sized large display. The global interaction space is used to issue commands, moving display output from on-demand selected computers to the large display and back again. It is also used to select multiple computers and concurrently execute commands on them.
Giacomo Tartari, Daniel Stødle, John Markus Bjørndalen, Phuong Hoai Ha, Otto J. Anshus
HSI5
2013 pVD - Personal Video Distribution
abstract
A user has several personal computers, including mobile phones, tablets, and laptops, and needs to watch live camera feeds from and videos stored at any of these computers at one or more of the others. Industry solutions designed for many users, computers, and videos can be complicated and slow to apply. The user must typically rely on a third party service or at least log in. The Personal Video Distribution (pVD) system supports sending and viewing live and stored videos between any of a single user's computers, and allows for a smooth handover of play back between computers. The system avoids any third parties, and relies only on the user's personal computers. We present the architecture, design and implementation of the pVD prototype. The architecture is comprised of functionality for sending videos, subscribing to videos, and maintaining the video play-back state. The design has a local side sending and viewing videos, and a global side coordinating the switching and distribution of videos, and maintaining subscriptions and video state. The prototype is primarily done in Python. A set of experiments was conducted to document the performance of the prototype. The results show that pVD global side has low CPU usage, and supports a handful of simultaneous exchanges of videos on a wireless network.
John Markus Bjørndalen, Phuong Hoai Ha, Otto J. Anshus
WiMob4
2011 A Step towards Making Local and Remote Desktop Applications Interoperable with High-Resolution Tiled Display Walls
Tor-Magne Stien Hagen, Daniel Stødle, John Markus Bjørndalen, Otto J. Anshus
DAIS4
2010 The Synchronization Power of Coalesced Memory Accesses
abstract
Multicore architectures have established themselves as the new generation of computer architectures. As part of the one core to many cores evolution, memory access mechanisms have advanced rapidly. Several new memory access mechanisms have been implemented in many modern commodity multicore architectures. By specifying how processing cores access shared memory, memory access mechanisms directly influence the synchronization capabilities of multicore architectures. Therefore, it is crucial to investigate the synchronization power of these new memory access mechanisms. This paper investigates the synchronization power of coalesced memory accesses, a family of memory access mechanisms introduced in recent large multicore architectures such as the Compute Unified Device Architecture (CUDA). We first define three memory access models to capture the fundamental features of the new memory access mechanisms. Subsequently, we prove the exact synchronization power of these models in terms of their consensus numbers. These tight results show that the coalesced memory access mechanisms can facilitate strong synchronization between the threads of multicore architectures, without the need of synchronization primitives other than reads and writes. In the case of the contemporary CUDA processors, our results imply that the coalesced memory access mechanisms have consensus numbers up to 64.
Phuong Hoai Ha, Philippas Tsigas, Otto J. Anshus
IEEE Trans. Parallel Distributed Syst.3
2009 NB-FEB: A Universal Scalable Easy-to-Use Synchronization Primitive for Manycore Architectures
Phuong Hoai Ha, Philippas Tsigas, Otto J. Anshus
OPODIS3
2009 Preliminary results on nb-feb, a synchronization primitive for parallel programming
abstract
We introduce a non-blocking full/empty bit primitive, or NB-FEB for short, as a promising synchronization primitive for parallel programming on may-core architectures. We show that the NB-FEB primitive is universal, scalable and feasible. NB-FEB, together with registers, can solve the consensus problem for an arbitrary number of processes (universality). NB-FEB is combinable, namely its memory requests to the same memory location can be combined into only one memory request, which consequently mitigates performance degradation due to synchronization "hot spots" (scalability). Since NB-FEB is a variant of the original full/empty bit that always returns a value instead of waiting for a conditional flag, it is as feasible as the original full/empty bit, which has been implemented in many computer systems (feasibility).
Phuong Hoai Ha, Philippas Tsigas, Otto J. Anshus
PPoPP3
2008 Liberating the Desktop
abstract
We report on a system supporting cross-platform mirroring of user-selectable regions from one or multiple computer desktops onto nearby network accessible projectors and displays (NADs). The purpose is a simple and flexible use of nearby display resources requiring no permanent installation of new software on the desktop computer. The NAD system architecture consists of a NAD side and a desktop side. The desktop software is downloaded to the desktop computer on demand, from a web server running on the NAD. The desktop and NAD software handle the integration of user-selectable desktop regions and remote control between the desktop computer and the NAD. The system is implemented in Java 1.6. At a resolution of 800 by 600 pixels the system supports mirroring of dynamic content at 38.6 fps. At 1600 by 1200 pixels the refresh rate is 12.85 fps. For static content such as images and slide show presentations the system's bandwidth usage is within the capacity of a 11 Mbit/s wireless network. For dynamic content such as videos and games the system requires at least a 100 Mbit/s connection.
Tor-Magne Stien Hagen, Espen Skjelnes Johnsen, Daniel Stødle, John Markus Bjørndalen, Otto J. Anshus
ACHI5
2008 Wait-free Programming for General Purpose Computations on Graphics Processors
abstract
The fact that graphics processors (GPUs) are today’s most powerful computational hardware for the dollar has motivated researchers to utilize the ubiquitous and powerful GPUs for general-purpose computing. Recent GPUs feature the single-program multiple-data (SPMD) multicore architecture instead of the single-instruction multiple-data (SIMD). However, unlike CPUs, GPUs devote their transistors mainly to data processing rather than data caching and flow control, and consequently most of the powerful GPUs with many cores do not support any synchronization mechanisms between their cores. This prevents GPUs from being deployed more widely for general-purpose computing. This paper aims at bridging the gap between the lack of synchronization mechanisms in recent GPU architectures and the need of synchronization mechanisms in parallel applications. Based on the intrinsic features of recent GPU architectures, we construct strong synchronization objects like wait-free and t-resilient read-modify-write objects for a general model of recent GPU architectures without strong hardware synchronization primitives like test-and-set and compare-and-swap. Accesses to the wait-free objects have time complexity O(N), whether N is the number of processes. Our result demonstrates that it is possible to construct wait-free synchronization mechanisms for GPUs without the need of strong synchronization primitives in hardware and that wait-free programming is possible for GPUs.
Phuong Hoai Ha, Philippas Tsigas, Otto J. Anshus
IPDPS3
2008 Wait-free programming for general purpose computations on graphics processors
abstract
This paper aims at bridging the gap between the lack of synchronization mechanisms in recent graphics processor (GPU) architectures and the need of synchronization mechanisms in parallel applications. Based on the intrinsic features of recent GPU architectures, we construct strong synchronization objects like wait-free and t-resilient read-modify-write objects for a general model of recent GPU architectures without strong hardware synchronization primitives like test-and-set and compare-and-swap. Accesses to the new wait-free objects have time complexity O(N), where N is the number of concurrent processes. The wait-free objects have space complexity O(N2), which is optimal. Our result demonstrates that it is possible to construct wait-free synchronization mechanisms for GPUs without the need of strong synchronization primitives in hardware and that wait-free programming is possible for GPUs.
Phuong Hoai Ha, Philippas Tsigas, Otto J. Anshus
PODC3
2008 The Synchronization Power of Coalesced Memory Accesses
Phuong Hoai Ha, Philippas Tsigas, Otto J. Anshus
DISC3
2005 Low Overhead High Performance Runtime Monitoring of Collective Communication
abstract
Scalability of parallel applications on clusters and multi-clusters is often limited by communication performance. Message tracing can provide data for understanding bottlenecks, and for performance tuning. However, it requires collecting, storing, analyzing, and transferring potentially gigabytes of data. We have designed the EventSpace system for low overhead and high performance runtime collective communication trace analysis. EventSpace separates the perturbation and performance requirements of data collection, analysis, gathering sand visualization. Data collection overhead is low since the minimum amount of data is recorded and stored temporarily in main memory. The recorded data is either discarded or analyzed on demand using available cluster resources. Analysis is distributed for high performance, and coscheduled with the computation and communication system threads for low perturbation. Gathering of analyzed data is done using extensible collective communication operations, which can be tuned to trade off between performance and monitoring overhead. EventSpace was used to do run-time monitoring and analysis of collective communication micro-benchmarks run on clusters, multi-clusters, and multi-clusters with emulated WAN links. Performance data was collected, analyzed and gathered with 0-3% monitoring overhead.
Lars Ailo Bongo, Otto J. Anshus, John Markus Bjørndalen
ICPP2
2004 Collective Communication Performance Analysis Within the Communication System
Lars Ailo Bongo, Otto J. Anshus, John Markus Bjørndalen
Euro-Par2
2003 EventSpace - Exposing and Observing Communication Behavior of Parallel Cluster Applications
Lars Ailo Bongo, Otto J. Anshus, John Markus Bjørndalen
Euro-Par2
2000 Operating system support for multimedia systems
Thomas Plagemann, Vera Goebel, Pål Halvorsen, Otto J. Anshus
Comput. Commun.4
1998 PastSet - An Efficient High Level Inter Process Communication Mechanism?
abstract
A new high level IPC mechanism, PastSet, is presented. PastSet supports partial causal ordered logging of synchronization and communication events. Communicated data are represented as tuples and stored in a common repository accessible by all processes. In the repository, tuples of the same template are causally ordered. While being equivalent to semaphores, messages and pipes, PastSet also offer a powerful abstraction suitable for knapsack type parallel applications as well as applications that require state logging. An implementation on uniprocessors and four-way multiprocessors running Linux. The mechanism is integrated into the Linux kernel alongside existing Sys V mechanisms. PastSet gives up to 40% faster synchronization and up to 56% better small package bandwidths than achieved with Linux semaphores, messages, and pipes. Although being a higher abstraction level mechanism, PastSet proves to consistently outperform existing Linux Sys V interprocess communication mechanisms on multiprocessors and to a large extend also on uniprocessors.
Brian Vinter, Otto J. Anshus, Tore Larsen
ICPP2
1996 Thread Scheduling for Cache Locality
abstract
This paper describes a method to improve the cache locality of sequential programs by scheduling fine-grained threads. The algorithm relies upon hints provided at the time of thread creation to determine a thread execution order likely to reduce cache misses. This technique may be particularly valuable when compiler-directed tiling is not feasible. Experiments with several application programs, on two systems with different cache structures, show that our thread scheduling method can improve program performance by reducing second-level cache misses.
James Philbin, Jan Edler, Otto J. Anshus, Craig C. Douglas, Kai Li 0001
ASPLOS3