EDBT 2026 Demo / reviewers in the wild / expert
Islam Atta
dblp:99/11514
· DBLP profile ↗
6ranked-venue papers
4as first author
0since 2021 · last 2015
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 3 first-authorSoftware engineering, systems software and programming languages · 2 · 1 first-authorDatabases, data management, data science and information retrieval · 2 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
4 papers |
Memory systems · 75% Processor architecture and microarchitecture · 13% Parallel and multicore computing · 6% | |
| Databases, data mining, and information retrieval
3 papers |
Transaction processing and concurrency control · 100% |
Topics — the 14 heaviest of 14, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Memory systems › cache › CPU cache
instruction cache |
0.5 | 3 | 2014 | ADDICT: Advanced Instruction Chasing for Transactions · Proc. VLDB Endow. 2014 STREX: boosting instruction cache reuse in OLTP workloads through stratified transaction execution · ISCA 2013 SLICC: Self-Assembly of Instruction Cache Collectives for OLTP Workloads · MICRO 2012 |
Memory systems
cache |
0.4 | 2 | 2015 | Self-contained, accurate precomputation prefetching · MICRO 2015 SLICC: Self-Assembly of Instruction Cache Collectives for OLTP Workloads · MICRO 2012 |
Memory systems › cache › prefetching
data prefetching |
0.2 | 1 | 2015 | Self-contained, accurate precomputation prefetching · MICRO 2015 |
Transaction processing and concurrency control
transaction scheduling |
0.2 | 1 | 2014 | ADDICT: Advanced Instruction Chasing for Transactions · Proc. VLDB Endow. 2014 |
Memory systems › memory hierarchy
cache hierarchy |
0.2 | 1 | 2014 | ADDICT: Advanced Instruction Chasing for Transactions · Proc. VLDB Endow. 2014 |
Memory systems
cache design |
0.2 | 1 | 2013 | STREX: boosting instruction cache reuse in OLTP workloads through stratified transaction execution · ISCA 2013 |
Memory systems › cache › cache behavior
cache reuse |
0.2 | 1 | 2013 | STREX: boosting instruction cache reuse in OLTP workloads through stratified transaction execution · ISCA 2013 |
Processor architecture and microarchitecture › instruction fetch
instruction prefetching |
0.1 | 1 | 2012 | SLICC: Self-Assembly of Instruction Cache Collectives for OLTP Workloads · MICRO 2012 |
Processor architecture and microarchitecture
multicore design |
0.1 | 1 | 2012 | SLICC: Self-Assembly of Instruction Cache Collectives for OLTP Workloads · MICRO 2012 |
Parallel and multicore computing › parallel programming runtimes › thread management
thread migration |
0.1 | 1 | 2012 | SLICC: Self-Assembly of Instruction Cache Collectives for OLTP Workloads · MICRO 2012 |
Transaction processing and concurrency control
OLTP |
0.1 | 2 | 2013 | STREX: boosting instruction cache reuse in OLTP workloads through stratified transaction execution · ISCA 2013 SLICC: Self-Assembly of Instruction Cache Collectives for OLTP Workloads · MICRO 2012 |
Memory systems › cache › prefetching
hardware prefetching |
0.1 | 1 | 2015 | Self-contained, accurate precomputation prefetching · MICRO 2015 |
Performance modeling and evaluation › workload characterization
memory characterization |
0.1 | 1 | 2014 | ADDICT: Advanced Instruction Chasing for Transactions · Proc. VLDB Endow. 2014 |
Performance modeling and evaluation
workload characterization |
0.1 | 1 | 2014 | ADDICT: Advanced Instruction Chasing for Transactions · Proc. VLDB Endow. 2014 |
Methods — techniques the papers use, named apart from their topics
thread migration · 0.6memory characterization · 0.4core assignment · 0.4instruction prefetching · 0.3program slicing · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2015 | Self-contained, accurate precomputation prefetchingabstractThis work revisits precomputation prefetching targeting long access latency loads with access patterns that are hard to predict. It presents Ekivolos, a precomputation prefetcher system that automatically builds prefetching slices that contain enough control flow instructions to faithfully and autonomously recreate the program's access behavior without inducing monitoring and execution overhead on the main thread. Ekivolos departs from the traditional notion of creating optimized short slices. In contrast, it shows that even longer slices can run ahead of the main thread and perform useful prefetches as long as they are sufficiently accurate. Ekivolos operates on arbitrary application binaries and takes advantage of the observed execution paths in creating its slices. On a set of emerging workloads Ekivolos outperforms three state-of-the-art hardware prefetchers and previously proposed precomputation-based prefetchers. Islam Atta, Xin Tong 0005, Vijayalakshmi Srinivasan, Ioana Baldini, Andreas Moshovos |
MICRO | 1 |
| 2014 | ADDICT: Advanced Instruction Chasing for TransactionsabstractRecent studies highlight that traditional transaction processing systems utilize the micro-architectural features of modern processors very poorly. L1 instruction cache and long-latency data misses dominate execution time. As a result, more than half of the execution cycles are wasted on memory stalls. Previous works on reducing stall time aim at improving locality through either hardware or software techniques. However, exploiting hardware resources based on the hints given by the software-side has not been widely studied for data management systems. In this paper, we observe that, independently of their high-level functionality, transactions running in parallel on a multicore system execute actions chosen from a limited sub-set of predefined database operations. Therefore, we initially perform a memory characterization study of modern transaction processing systems using standardized benchmarks. The analysis demonstrates that same-type transactions exhibit at most 6% overlap in their data footprints whereas there is up to 98% overlap in instructions. Based on the findings, we design ADDICT, a transaction scheduling mechanism that aims at maximizing the instruction cache locality. ADDICT determines the most frequent actions of database operations, whose instruction footprint can fit in an L1 instruction cache, and assigns a core to execute each of these actions. Then, it schedules each action on its corresponding core. Our prototype implementation of ADDICT reduces L1 instruction misses by 85% and the long latency data misses by 20%. As a result, ADDICT leads up to a 50% reduction in the total execution time for the evaluated workloads. Pinar Tözün, Islam Atta, Anastasia Ailamaki, Andreas Moshovos |
Proc. VLDB Endow. | 2 |
| 2013 | A dual grain hit-miss detector for large die-stacked DRAM cachesabstractDie-Stacked DRAM caches offer the promise of improved performance and reduced energy by capturing a larger fraction of an application's working set than on-die SRAM caches. However, given that their latency is only 50% lower than that of main memory, DRAM caches considerably increase latency for misses. They also incur a significant energy overhead for remote lookups in snoop-based multi-socket systems. Ideally, it would be possible to detect in advance that a request will miss in the DRAM cache and thus selectively bypass it. This work proposes a “dual grain filter” which successfully predicts whether an access is a hit or a miss in most cases. Experimental results with commercial and scientific workloads show that a 158KB dual-grain filter can correctly predict data block residency for 85% of all accesses to a 256MB DRAM cache. As a result, average off-die latency with our filter is within 8% of that possible with a perfectly accurate filter, which is impractical to implement. Michel El-Nacouzi, Islam Atta, Misel-Myrto Papadopoulou, Jason Zebchuk, Natalie D. Enright Jerger, Andreas Moshovos |
DATE | 2 |
| 2013 | STREX: boosting instruction cache reuse in OLTP workloads through stratified transaction executionabstractOnline transaction processing (OLTP) workload performance suffers from instruction stalls; the instruction footprint of a typical transaction exceeds by far the capacity of an L1 cache, leading to ongoing cache thrashing. Several proposed techniques remove some instruction stalls in exchange for error-prone instrumentation to the code base, or a sharp increase in the L1-I cache unit area and power. Others reduce instruction miss latency by better utilizing a shared L2 cache. SLICC [2], a recently proposed thread migration technique that exploits transaction instruction locality, is promising for high core counts but performs sub-optimally or may hurt performance when running on few cores. Islam Atta, Pinar Tözün, Xin Tong 0005, Anastasia Ailamaki, Andreas Moshovos |
ISCA | 1 |
| 2012 | Reducing OLTP instruction misses with thread migrationabstractDuring an instruction miss a processor is unable to fetch instructions. The more frequent instruction misses are the less able a modern processor is to find useful work to do and thus performance suffers. Online transaction processing (OLTP) suffers from high instruction miss rates since the instruction footprint of OLTP transactions does not fit in today's L1-I caches. However, modern many-core chips have ample aggregate L1 cache capacity across multiple cores. Looking at the code paths concurrently executing transactions follow, we observe a high degree of repetition both within and across transactions. This work presents TMi a technique that uses thread migration to reduce instruction misses by spreading the footprint of a transaction over multiple L1 caches. TMi is a software-transparent, hardware technique; TMi requires no code instrumentation, and efficiently utilizes available cache capacity. This work evaluates TMi's potential and shows that it may reduce instruction misses by 51% on average. This work discusses the underlying tradeoffs and challenges, such as an increase in data misses, and points to potential solutions. Islam Atta, Pinar Tözün, Anastasia Ailamaki, Andreas Moshovos |
DaMoN | 1 |
| 2012 | SLICC: Self-Assembly of Instruction Cache Collectives for OLTP WorkloadsabstractOnline transaction processing (OLTP) is at the core of many data center applications. OLTP workloads are known to have large instruction footprints that foil existing L1 instruction caches resulting in poor overall performance. Prefetching can reduce the impact of such instruction cache miss stalls, however, state-of-the-art solutions require large dedicated hardware tables on the order of 40KB in size. SLICC is a programmer transparent, low cost technique to minimize instruction cache misses when executing OLTP workloads. SLICC migrates threads, spreading their instruction footprint over several L1 caches. It exploits repetition within and across transactions, where a transaction's first iteration prefetches the instructions for subsequent iterations or similar subsequent transactions. SLICC reduces instruction misses by 58% on average for TPC-C and TPCE, thereby improving performance by 68%. When compared to a state-of-the-art prefetcher, and notwithstanding the increased storage overheads (42x as compared to SLICC), performance using SLICC is 21% higher for TPC-E and within 2% for TPC-C. Islam Atta, Pinar Tözün, Anastasia Ailamaki, Andreas Moshovos |
MICRO | 1 |