VLDB 2026 Research / reviewers in the wild / expert
Rachid Karami
dblp:348/7628
· DBLP profile ↗
5ranked-venue papers
2as first author
5since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 4 · 2 first-author · 4 since 2021Systems, architecture and hardware · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | RTG: Exploring the Dynamic Scheduling Space of Real-Time Generative AI workloads on Emerging Heterogeneous SystemsabstractThe integration of Large Language Model (LLM) agents into interactive applications—driven by demands for enhanced productivity, multitasking, and personalization—has shifted generative AI toward on-device execution to ensure data privacy and availability. This trend gives rise to Real-Time Generative (RTG) workloads, in which compute-intensive LLM inference runs concurrently with latency-sensitive real-time tasks on heterogeneous SoCs integrating CPU, GPU, and NPU backends. However, significant variability across LLM stages (prefill vs. decode) and input-dependent sequence and context lengths creates a complex and dynamic scheduling space that conventional policies fail to handle. In this work, we define realistic RTG scenarios and evaluate six scheduling policies. Our results show that an RTG-aware scheduler adapts to hardware-specific preferences and stage-level divergence, achieving near-standalone (dedicated-resource) LLM performance at a 1.83% deadline violation rate. These results highlight the need for workload-aware, dynamic scheduling to enable on-device RTG applications. Rachid Karami, Rajeev Patwari, Hyoukjun Kwon, Ashish Sirasao |
ISPASS | 1 |
| 2026 | Characterizing State Space Model and Hybrid Language Model Performance with Long ContextabstractEmerging applications such as AR are driving demands for machine intelligence capable of processing continuous and/or long-context inputs on local devices. However, currently dominant models based on Transformer architecture suffers from the quadratic computational and memory overhead, which hinders applications required to process long contexts. This has spurred a paradigm shift towards new architectures like State Space Models (SSMs) and SSM-Transformer hybrid models, which provide near-linear scaling. The near-linear scaling enabled efficient handling of millions of tokens while delivering high performance in recent studies. Although such works present promising results, their workload characteristics in terms of computational performance and hardware resource requirements are not yet thoroughly explored, which limits our understanding of their implications to the system level optimizations. To address this gap, we present a comprehensive, comparative benchmarking of carefully selected Transformers, SSMs, and hybrid models specifically for long-context inference on consumer and embedded GPUs. Our analysis shows that SSMs are well-suited for on-device AI on consumer and embedded GPUs for long context inferences. While Transformers are up to $1.9 \times$ faster at short sequences ($ \lt 8 \mathrm{~K}$ tokens), SSMs demonstrate a dramatic performance inversion, becoming up to $4 \times$ faster at very long contexts ($\sim 57 \mathrm{~K}$ tokens), thanks to their linear computational complexity and $\boldsymbol{\sim} \mathbf{6 4 \%}$ reduced memory footprint. Our operator-level analysis reveals that custom SSM kernels like selective scan despite being hardware-aware to minimize memory IO, dominate the inference runtime on edge platforms, accounting for over $55 \%$ of latency due to their sequential, element-wise nature. To foster further research, we are sharing CPU/GPU profiling traces and have made our characterization framework SSM-Scope open-sourced at https://github.com/sapmitra/ssm-scope Saptarshi Mitra, Rachid Karami, Haocheng Xu, Sitao Huang, Hyoukjun Kwon |
ISPASS | 2 |
| 2025 | Understanding the Performance Horizon of the Latest ML Workloads with NonGEMM WorkloadsabstractAmong ML operators today, GEneralMatrix Multiplication (GEMM)-based operators are known to be key operators that build the main backbone of ML models. As their computational overhead dominates the overall execution time (e.g., 42.8%-96.6% in our results), GEMM operators have been the prime optimization targets for fast ML inference. This led to advanced GPUs and accelerators available today, which provided significant boost in the GEMM performance compared to CPUs, aligned with the lesson from Amdahl's law. However, accelerating GEMM has significantly shifted the Amdahl's law's landscape for ML inference; due to the decreased GEMM execution time, the relative execution time of non-GEMM operators is now significant. Although the importance of non-GEMM performance is increasing, we have little knowledge about the non-GEMM performance horizon in the latest hardware platforms and models. Therefore, to guide non-GEMM-oriented optimizations, we conduct a thorough performance analysis of 17 widely adopted ML models in Hugging Face and Torchvision on workstation and data center platforms with/without GPUs. We discover that non-GEMM performance bottleneck is a considerable issue across all the platforms and models, accounting for 11.3% to 73.6% of total latency, on average. The challenge significantly aggravates when we apply quantization, which is a common model compression technique, due to the boosted GEMM performance and extra non-GEMM operators for dequantization and requantization. To provide insights into non-GEMM optimization targets, we demystify the most dominant non-GEMM operators for each model and deployment software. We also show that widely adopted optimizations such as operator fusion do not completely address the non-GEMM performance bottleneck, where nonGEMM operators still account for 15% to 48% of total latency. We will open-source our non-GEMM-oriented benchmark framework to facilitate research in non-GEMM optimization. Rachid Karami, Sheng-Chun Kao, Hyoukjun Kwon |
ISPASS | 1 |
| 2023 | Information Processing Factory 2.0 - Self-awareness for Autonomous Collaborative SystemsabstractThis paper summarizes the talks of a special session on the IPF 2.0 project, a collaborative German-US research project that leverages self-awareness principles for the self-management of distributed systems of autonomous multiprocessor systems-on-chip (MPSoCs). Nora Sperling, Alex Bendrick, Dominik Stöhrmann, Rolf Ernst, Bryan Donyanavard, Florian Maurer 0003, Oliver Lenke, Anmol Surhonne, Andreas Herkersdorf, Walaa Amer, Caio Batista de Melo, Ping-Xiang Chen, Quang Anh Hoang, Rachid Karami, Biswadip Maity, Paul Nikolian, Mariam Rakka, Dongjoo Seo, Saehanseul Yi, Minjun Seo, Nikil Dutt, Fadi J. Kurdahi |
DATE | 14 |
| 2023 | Hardware Implementation and Evaluation of an Information Processing FactoryabstractThe Information Processing Factory (IPF) utilizes factory management principles to tackle the complexities of integrated embedded systems, ensuring continuous safe operation and optimization at runtime. This paper presents a hardware implementation of IPF that enables dynamic task migration across system resources, ensuring reliability in the face of internal or external failures. We demonstrate the effectiveness of IPF through the efficient migration of tasks in multiprocessor SoCs using a safety-critical pacemaker application as a case study. Despite the additional software and hardware requirements, implementing IPF in a pacemaker results in comparable reliability to dual modular redundancy (DMR) with faster service resumption and improved resource utilization. Walaa Amer, Mariam Rakka, Rachid Karami, Minjun Seo, Mazen A. R. Saghir, Rouwaida Kanj, Fadi J. Kurdahi |
VLSI-SoC | 3 |