Alfred Bratterud

dblp:142/7811 · DBLP profile ↗
← Back
6ranked-venue papers
3as first author
1since 2021 · last 2026
0000-0002-7342-191XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Computer networks · 1Theory of computation · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
Memory systems · 61% Storage systems · 30% Cloud and datacenter computing · 9%

Topics — the 4 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Memory systems › memory disaggregation
CXL memory
1.012026
No Atomics, No Problem. Developing a RAG Pipeline for Shared CXL Memory · IEEE Trans. Computers 2026
Storage systems
file systems
1.012026
No Atomics, No Problem. Developing a RAG Pipeline for Shared CXL Memory · IEEE Trans. Computers 2026
Memory systems
shared memory
1.012026
No Atomics, No Problem. Developing a RAG Pipeline for Shared CXL Memory · IEEE Trans. Computers 2026
Cloud and datacenter computing › cluster resource management and scheduling
cluster resource management
0.312026
No Atomics, No Problem. Developing a RAG Pipeline for Shared CXL Memory · IEEE Trans. Computers 2026

Methods — techniques the papers use, named apart from their topics

vector similarity search · 1.0retrieval-augmented generation · 1.0
YearPublicationVenuePosition
2026 No Atomics, No Problem. Developing a RAG Pipeline for Shared CXL Memory
abstract
We share early experiments with software development for shared CXL memory on the H3 Falcon C5022 CXL switch. We describe how the different stages in the development of a commercial RAG pipeline were impacted by the presence of CXL memory. All of Wikipedia is split into 46 million text passages which are embedded into a high-dimensional embedding space and subjected to heavy load in single-and multi-host experiments. We observe significant performance advantages available to DuckDB and Faiss without changing the software; we measure up to 12× latency reduction with real workloads and 40× reduction with synthetic workloads on CXL, compared to the same queries run with demand paging on NVMe. We explain the current challenges of working with two disjoint cache coherency domains and explain how the Fabric-Attached Memory File System (famfs) and famfs producer-consumer queues provide synchronization patterns without atomic operations on current ×86 CPUs. We then discuss open challenges remaining for high performance atomics and locking mechanisms in shared memory. Lastly we show how famfs page-level interleaving enables near-linear throughput scaling when a second host serves queries from the same Faiss index on shared fabric-attached memory and how shared memory allocations can be orchestrated by Kubernetes in a full vertical RAG deployment.
Alfred Bratterud, Gisle Dankel, Amin Farajianzadeh, John Groves, Chengyi Juan, Joshua Suetterlein, Andrés Márquez 0001, Petter Gustad, Arnt Emil Ingulstad
IEEE Trans. Computers1
2018 A Queue Model for Reliable Forecasting of Future CPU Consumption
Hugo Hammer, Anis Yazidi, Alfred Bratterud, Hårek Haugerud, Boning Feng
Mob. Networks Appl.3
2017 Unikernels for Cloud Architectures: How Single Responsibility can Reduce Complexity, Thus Improving Enterprise Cloud Security
abstract
ACKNOWLEDGEMENTS This work was in part funded by the European Commission through grant agreement no 644962 (PRISMACLOUD).
Andreas Happe, Bob Duncan, Alfred Bratterud
COMPLEXIS3
2017 The concept of workload delay as a quality-of-service metric for consolidated cloud environments with deadline requirements
abstract
Virtual Machine (VM) consolidation in the cloud has received significant research interest. A large body of approaches for VM consolidation in data centers resort to variants of the bin packing problem which tries to minimize the number of deployed physical machines while meeting the Service-Level-Agreement (SLA) constraints. In this paper we introduce the concept of workload delay as a Quality-of-Service (QoS) metric that captures directly the resulting degradation that a cloud user would experience in the case where the SLA is violated. Our results, that are based on real-life trace-based simulations, show that consolidating VMs based on the level of utilization results in little control over the resulting delay, a particularly significant drawback when running jobs with deadline requirements, while we are able to control the delay much better if we take into account our suggested metric of the delay.
Evangelos Tasoulas 0001, Hugo Hammer, Hårek Haugerud, Anis Yazidi, Alfred Bratterud, Boning Feng
NCA5
2015 IncludeOS: A Minimal, Resource Efficient Unikernel for Cloud Services
abstract
The emergence of cloud computing as a ubiquitous platform for elastically scaling services has generated need and opportunity for new types of operating systems. A service that needs to be both elastic and resource efficient needs A) highly specialized components, and B) to run with minimal resource overhead. Classical general purpose operating systems designed for extensive hardware support are by design far from meeting these requirements. In this paper we present IncludeOS, a single tasking library operating system for cloud services, written from scratch in C++. Key features include: extremely small disk-and memory footprint, efficient asynchronous I/O, OS-library where only what your service needs gets included, and only one device driver by default (virtio). As a test case a bootable disk image consisting of a simple DNS server with OS included is shown to require only 158 kb of disk space and to require 5-20% less CPU-time, depending on hardware, compared to the same binary running on Linux.
Alfred Bratterud, Alf-Andre Walla, Hårek Haugerud, Paal E. Engelstad, Kyrre M. Begnum
CloudCom1
2013 Maximizing Hypervisor Scalability Using Minimal Virtual Machines
abstract
The smallest instance offered by Amazon EC2 comes with 615MB memory and a 7.9GB disk image. While small by today's standards, embedded web servers with memory footprints well under 100kB, indicate that there is much to be saved. In this work we investigate how large VM-populations the open Stack hyper visor can be made to sustain, by tuning it for scalability and minimizing virtual machine images. Request-driven Qemu images of 512 byte are written in assembly, and more than 110 000 such instances are successfully booted on a 48 core host, before memory is exhausted. Other factors are shown to dramatically improve scalability, to the point where 10 000 virtual machines consume no more than 2.06% of the hyper visor CPU.
Alfred Bratterud, Hårek Haugerud
CloudCom (1)1