Michael Dang'ana

dblp:254/3540 · DBLP profile ↗
← Back
3ranked-venue papers
3as first author
3since 2021 · last 2026
0009-0003-2736-1113ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 3 first-author · 3 since 2021
YearPublicationVenuePosition
2026 FractalSort: High Precision Compressed Radix Sort on FPGA
abstract
State-of-the-art large data set high-precision sorting algorithms typically use hardware-accelerated radix sort. Advances in Dynamic Random Access Memory, Flash and High Bandwidth Memory (HBM) have enabled faster bandwidth intensive merge operations, where distribution-dependent data pre-processing techniques such as stochastic sampling bucketing offer alternatives impacted by increased data passes and vulnerability to data skew.This work addresses these limitations by introducing a compressed radix-sorting scheme for high-precision keys. Whereas radix sort histograms grow exponentially with precision, FractalSort guarantees bounded histogram size through the novel compression scheme which translates into smaller sorting circuits and reduced memory usage. Another key contribution is the novel optimized merge algorithm, which eliminates the need for data pre-processing and bucketing leading to higher bandwidth efficiency and reduced algorithm complexity. Using a tree-based recursive sorting architecture for space efficiency and low latency, the algorithm achieves fast sorting that exceeds the state-of-the-art on the CPU, FPGA, and GPU by 6x, 2.5x, and 3x bandwidth-adjusted throughput on 4GB to 2TB data sets. FractalSort is implemented on Xilinx Virtex UltraScale+ FPGA empirically demonstrating on-chip sorting at 20 Tb/s capable of memory-to-memory sorting of 32-bit keys at 3.2Tb/s using HBM.
Michael Dang'ana, Hans-Arno Jacobsen
IEEE Trans. Computers1
2026 Ksurf-Drone: Attention Kalman Filter for Contextual Bandit Optimization in Cloud Resource Allocation
abstract
Resource orchestration and configuration parameter search are key concerns for container-based infrastructure in cloud data centers. Large configuration search space and cloud uncertainties are often mitigated using contextual bandit techniques for resource orchestration including the state-of-the-art Drone orchestrator. Complexity in the cloud provider environment due to varying numbers of virtual machines introduces variability in workloads and resource metrics, making orchestration decisions less accurate due to increased nonlinearity and noise. Ksurf, a state-of-the-art variance-minimizing estimator method ideal for highly variable cloud data, enables optimal resource estimation under conditions of high cloud variability. This work evaluates the performance of Ksurf on estimation-based resource orchestration tasks involving highly variable workloads when employed as a contextual multi-armed bandit objective function model for cloud scenarios using Drone. Ksurf enables significantly lower latency variance of over$40\%$at p95 and p99, demonstrates significant reduction in CPU and master node memory usage on Kubernetes, resulting in a$7\%$cost savings in average worker pod count on$VarBench$Kubernetes benchmark.
Michael Dang'ana, Yuqiu Zhang, Hans-Arno Jacobsen
IEEE Trans. Cloud Comput.1
2025 Ksurf+: Attention Kalman Filter for Prediction Under Highly Variable Cloud Workloads
abstract
Resource estimation and workload forecasting are critical in cloud data centers. Complexity in the cloud provider environment due to varying numbers of virtual machines introduces high variability in workloads and resource usage, making estimations problematic using state-of-the-art models that fail to deal with nonlinear characteristics. High measurement noise and variance affect the estimation of resource metrics of cloud systems across packet networks influenced by unknown external dynamics. An ideal solution to these problems is the Kalman filter, a variance-minimizing estimator, ideal for highly variable data. This work introduces Ksurf+, a novel Kalman filter estimator using selective principal component analysis and an attention mechanism for enhanced short-horizon prediction. Ksurf+ improves prediction accuracy by 37% over state-of-the-art Kalman filters in prediction tasks, reduces the time series prediction error of the state-of-the-art Bi-directional Grid Long Short-Term Memory neural network by over 40%, improves Kafka workload-based scaling stability by 58%, reduces Kafka queue size and lowers Kubernetes worker pod CPU usage by 11.6% on the$VarBench$benchmark.
Michael Dang'ana, Hans-Arno Jacobsen
IEEE Trans. Cloud Comput.1