VLDB 2026 Research / reviewers in the wild / expert
Iris Uwizeyimana
dblp:332/1669
· DBLP profile ↗
4ranked-venue papers
2as first author
4since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 3 · 2 first-author · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CPU Bottlenecks in Rack-Stand Brain-Computer Interfaces: A Phase-Aware CharacterisationabstractRack-stand brain-computer interfaces (BCIs) remain the dominant platform for exploratory and clinical BCI research, yet their CPU behaviour is poorly understood [1]–[7]. We characterise four rack-stand BCI pipelines—seizure detection [8], movement intent [9], speech decoding [10], and neural-data compression [11]—using peer-reviewed implementations and redistributable datasets on a commodity x86 CPU. Using offline replay of recorded neural streams, we combine end-to-end runtime across three core-frequency operating points with phase-aware Top-Down execution-time and memory-boundedness breakdowns [12]. The results show that rack-stand BCIs do not form a single bottleneck category, that channel count can change CPU demand by about $3 \times$, and that a single “memory-bound” label is too coarse because different cache and memory levels dominate in different stages. Victor Kariofillis, Iris Uwizeyimana, Sara Ahmad, Tanvi Manku, Raghavendra Pradyumna Pothukuchi, Abhishek Bhattacharjee, Natalie D. Enright Jerger |
ISPASS | 2 |
| 2026 | A Heterogeneous Mapping of a Speech Neuroprosthesis ApplicationabstractHeterogeneous processors that combine CPUs, GPUs, and NPUs have become an important response to dark silicon and rising workload demands. The high performance per watt potential of these platforms offers a path to lower the energy consumption of ML applications, which is critical for the sustainable growth of these workloads. However, realizing these benefits requires understanding how different phases of an application interact with the compute and memory behavior of each hardware component. In this work, we use a machine-learning-based speech neuroprosthesis application to demonstrate how heterogeneous hardware can be leveraged to improve end-to-end performance. We introduce techniques for reducing memory usage, scheduling data movement, and managing synchronization, which enable different stages of the application to execute efficiently across the NPU, GPU, and multi-threaded CPU. Our evaluation shows a clear divergence in hardware preferences when optimizing for latency versus throughput per watt, and highlights the role of interference when deploying the application in a pipeline-parallel manner. This work shows that understanding both the application structure and the underlying hardware is essential for effectively exploiting heterogeneity and offers guidance for designing sustainable heterogeneous application mappings. Iris Uwizeyimana, Oscar Sun, Natalie D. Enright Jerger |
ISPASS | 1 |
| 2025 | Carbon-Aware Server ReplacementabstractCloud computing contributes significantly to the global carbon emissions. Server management, including server replacement policies heavily influences the carbon footprint of cloud computing. While periodically replacing servers with newer hardware enhances energy efficiency, in turn reducing the operational carbon emissions from running servers, it also increases the embodied carbon footprint associated with server manufacturing. We propose an analytical model to determine the optimal server replacement frequency to reduce the overall datacenter carbon footprint, encompassing both embodied and operational carbon. Using our model, we perform a case study to analyze the impact prolonged server lifetime has on the net carbon footprint of a high performance computing cluster. Iris Uwizeyimana, Natalie D. Enright Jerger |
ISPASS | 1 |
| 2022 | ALTOCUMULUS: Scalable Scheduling for Nanosecond-Scale Remote Procedure CallsabstractOnline services in modern datacenters use Remote Procedure Calls (RPCs) to communicate between different software layers. Despite RPCs using just a few small functions, inefficient RPC handling can cause delays to propagate across the system and degrade end-to-end performance. Prior work has reduced RPC processing time to less than 1 $\mu$ s, which now shifts the bottleneck to the scheduling of RPCs. Existing RPC schedulers suffer from either high overheads, inability to effectively utilize high core-count CPUs or do not adaptively fit different traffic patterns. To address these shortcomings, we present ALTOCUMULUS,1a scalable, software-hardware codesign to schedule RPCs at nanosecond scales. ALTOCUMULUS provides a proactive scheduling scheme and low-overhead messaging mechanism on top of a decentralized user runtime. ALTOCUMULUS also offers direct access from the user space to a set of simple hardware primitives to quickly migrate long-latency RPCs. We evaluate ALTOCUMULUS with synthetic workloads and an end-to-end in-memory key-value store application under real-world traffic patterns. ALTOCUMULUS improves throughput by 1.3-24.6$\times$ under a 99thpercentile latencythpercentile latency $\lt 8.5\mu \mathrm{s}$.1Automatic Concurrent Migration Load-balancing Strategy (AutoCuMuLuS), homophonic with “altocumulus” as a type of clouds in meteorology, fragmented to separate patches or nodes. Jiechen Zhao 0002, Iris Uwizeyimana, Karthik Ganesan 0002, Mark C. Jeffrey, Natalie D. Enright Jerger |
MICRO | 2 |