Saba Jamilan

dblp:199/3535 · DBLP profile ↗
← Back
4ranked-venue papers
3as first author
3since 2021 · last 2026
0000-0003-2259-6426ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 RIFS: Run-Time Invariant Function Specialization
abstract
Compilers apply optimizations such as function specialization and constant propagation to eliminate redundant work at compile time. However, because compilers must prove that values are constant, many profitable optimization opportunities remain unrealized. In this paper, we propose run-time invariant function specialization (RIFS), a profile-guided compiler technique that specializes functions based on runtime invariant call-site-specific argument values. RIFS introduces a novel value-profiling LLVM pass to identify runtime invariant arguments, even though they cannot be proven constant statically. A subsequent LLVM transformation pass generates specialized function variants tailored to these value profiles. To efficiently select among potentially thousands of specialization candidates, we develop a predictive cost model that estimates the performance benefit of each candidate prior to code generation. We integrate our passes seamlessly into the existing PGO-enabled LLVM toolchain. We evaluate RIFS across 11 real-world applications, demonstrating substantial improvements over state-of-the-art optimization techniques. RIFS achieves an average speedup of 6.3% and an instruction reduction of 2.5% over the LLVM -O3+PGO baseline.
Saba Jamilan, Snehasish Kumar, Heiner Litz
CC1
2024 RPG2: Robust Profile-Guided Runtime Prefetch Generation
abstract
Data cache prefetching is a well-established optimization to overcome the limits of the cache hierarchy and keep the processor pipeline fed with data. In principle, accurate, well-timed prefetches can sidestep the majority of cache misses and dramatically improve performance. In practice, however, it is challenging to identify which data to prefetch and when to do so. In particular, data can be easily requested too early, causing eviction of useful data from the cache, or requested too late, failing to avoid cache misses. Competition for limited off-chip memory bandwidth must also be balanced between prefetches and a program's regular "demand" accesses. Due to these challenges, prefetching can both help and hurt performance, and the outcome can depend on program structure, decisions about what to prefetch and when to do it, and, as we demonstrate in a series of experiments, program input, processor microarchitecture, and their interaction as well.
Nathan Sobotka, Soyoon Park, Saba Jamilan, Tanvir Ahmed Khan 0001, Baris Kasikci, Gilles Pokam, Heiner Litz, Joseph Devietti
ASPLOS (2)4
2022 APT-GET: profile-guided timely software prefetching
abstract
Prefetching which predicts future memory accesses and preloads them from main memory, is a widely-adopted technique to overcome the processor-memory performance gap. Unfortunately, hardware prefetchers implemented in today's processors cannot identify complex and irregular memory access patterns exhibited by modern data-driven applications and hence developers need to rely on software prefetching techniques. We investigate the challenges of enabling effective, automated software data prefetching. Our investigation reveals that the state-of-the-art compiler-based prefetching mechanism falls short in achieving high performance due to its static nature. Based on this insight, we design APT-GET, a novel profile-guided technique that ensures prefetch timeliness by leveraging dynamic execution time information. APT-GET leverages efficient hardware support such as Intel's Last Branch Record (LBR), for collecting application execution profiles with negligible overhead to characterize the execution time of loads. APT-GET then introduces a novel analytical model to find the optimal prefetch-distance and prefetch injection site based on the collected profile to enable timely prefetches. We study APT-GET in the context of 10 real-world applications and demonstrate that it achieves a speedup of up to 1.98× and of 1.30× on average. By ensuring prefetch timeliness, APT-GET improves the performance by 1.25× over the state-of-the-art software data prefetching mechanism.
Saba Jamilan, Tanvir Ahmed Khan 0001, Grant Ayers, Baris Kasikci, Heiner Litz
EuroSys1
2017 Cache Energy Management through Dynamic Reconfiguration Approach in Opto-Electrical NoC
abstract
Multi/Many-core architectures will be the popular platform for future system design. Recent investigations show that the hybrid optical-electrical interconnection network can be an appropriate alternative to the traditional electrical NoC. Undoubtedly, memory wall is one of the most important challenges of multi/many-core systems which can somehow be alleviated thanks to hierarchical memory structure. Cache subsystem plays an essential role in increasing the efficiency of the memory structure. In this paper, after exploring the effect of cache subsystem's parameters in a many-core platform with opto-electrical interconnect, we propose a mechanism to increase the cache energy efficiency. The simulation results of the proposed approach show that after applying our method, the energy efficiency parameter improves by 17% and 23% in SPLASH2 and PARSEC benchmarks, respectively.
Saba Jamilan, Meisam Abdollahi, Siamak Mohammadi
PDP1