EDBT 2026 Demo / reviewers in the wild / expert
Anahita Shayesteh
dblp:40/5388
· DBLP profile ↗
11ranked-venue papers
2as first author
0since 2021 · last 2020
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 9 · 2 first-authorSoftware engineering, systems software and programming languages · 4
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
3 papers |
Storage systems · 49% Performance modeling and evaluation · 25% GPUs and heterogeneous computing · 13% |
Topics — the 9 heaviest of 10, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Storage systems
flash and SSD |
0.5 | 2 | 2018 | Performance Characterization of NVMe-over-Fabrics Storage Disaggregation · ACM Trans. Storage 2018 Performance Characterization of Hyperscale Applicationson on NVMe SSDs · SIGMETRICS 2015 |
Performance modeling and evaluation
workload characterization |
0.5 | 2 | 2018 | Performance Characterization of NVMe-over-Fabrics Storage Disaggregation · ACM Trans. Storage 2018 Performance Characterization of Hyperscale Applicationson on NVMe SSDs · SIGMETRICS 2015 |
Storage systems › distributed storage
disaggregated storage |
0.3 | 1 | 2018 | Performance Characterization of NVMe-over-Fabrics Storage Disaggregation · ACM Trans. Storage 2018 |
Storage systems › networked storage › storage networking
NVMe over Fabrics |
0.3 | 1 | 2018 | Performance Characterization of NVMe-over-Fabrics Storage Disaggregation · ACM Trans. Storage 2018 |
GPUs and heterogeneous computing
control flow divergence |
0.2 | 1 | 2013 | SIMD divergence optimization through intra-warp compaction · ISCA 2013 |
Processor architecture and microarchitecture
SIMD |
0.2 | 1 | 2013 | SIMD divergence optimization through intra-warp compaction · ISCA 2013 |
Cloud and datacenter computing
datacenter storage |
0.1 | 1 | 2018 | Performance Characterization of NVMe-over-Fabrics Storage Disaggregation · ACM Trans. Storage 2018 |
Performance modeling and evaluation
storage performance evaluation |
0.1 | 1 | 2015 | Performance Characterization of Hyperscale Applicationson on NVMe SSDs · SIGMETRICS 2015 |
Parallel and multicore computing › data parallelism
data-parallel applications |
0.0 | 1 | 2013 | SIMD divergence optimization through intra-warp compaction · ISCA 2013 |
Methods — techniques the papers use, named apart from their topics
stress testing · 0.3performance measurement · 0.3benchmarking · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2020 | Modeling Analytics for Computational StorageabstractNext generation flash storage will be armed with a substantial amount of computing power. In this paper, we investigate opportunities to utilize this computational capability to optimize Online Analytical Processing (OLAP) applications. We have directed our analysis at the performance of a subset of TPC-DS queries using Hadoop clusters and two database engines, SPARK-SQL and Presto. We model the expected speed-up achieved by offloading a few operations that are executed first within most SQL plans. Offloading these operations requires minimal cooperation from the database engine, and no changes to the existing plan. We show that the speed-up achieved varies significantly among queries and between engines, and that the queries benefiting the most are I/O heavy with high selectivity of the "needle in the haystack" variety. Our main contribution is estimating the speed-up anticipated from pushing the execution of a few key SQL building blocks (scan, filter, and project operations) to computational storage when using read optimized, columnar Parquet format files. Veronica Lagrange Moutinho dos Reis, Harry (Huan) Li, Anahita Shayesteh |
ICPE | 3 |
| 2018 | Performance Characterization of NVMe-over-Fabrics Storage DisaggregationabstractStorage disaggregation separates compute and storage to different nodes to allow for independent resource scaling and, thus, better hardware resource utilization. While disaggregation of hard-drives storage is a common practice, NVMe-SSD (i.e., PCIe-based SSD) disaggregation is considered more challenging. This is because SSDs are significantly faster than hard drives, so the latency overheads (due to both network and CPU processing) as well as the extra compute cycles needed for the offloading stack become much more pronounced. In this work, we characterize the overheads of NVMe-SSD disaggregation. We show that NVMe-over-Fabrics (NVMe-oF)—a recently released remote storage protocol specification—reduces the overheads of remote access to a bare minimum, thus greatly increasing the cost-efficiency of Flash disaggregation. Specifically, while recent work showed that SSD storage disaggregation via iSCSI degrades application-level throughput by 20%, we report on negligible performance degradation with NVMe-oF—both when using stress-tests as well as with a more-realistic KV-store workload. Zvika Guz, Harry (Huan) Li, Anahita Shayesteh, Vijay Balakrishnan |
ACM Trans. Storage | 3 |
| 2017 | NVMe-over-fabrics performance characterization and the path to low-overhead flash disaggregationabstractStorage disaggregation separates compute and storage to different nodes in order to allow for independent resource scaling and thus, better hardware resource utilization. While disaggregation of hard-drives storage is a common practice, NVMe-SSD (i.e., PCIe-based SSD) disaggregation is considered more challenging. This is because SSDs are significantly faster than hard drives, so the latency overheads (due to both network and CPU processing) as well as the extra compute cycles needed for the offloading stack become much more pronounced. Zvika Guz, Harry (Huan) Li, Anahita Shayesteh, Vijay Balakrishnan |
SYSTOR | 3 |
| 2015 | Performance Characterization of Hyperscale Applicationson on NVMe SSDsabstractThe storage subsystem has undergone tremendous innovation in order to keep up with the ever-increasing demand for throughput. NVMe based SSDs are the latest development in this domain, delivering unprecedented performance in terms of both latency and peak bandwidth. Given their superior performance, NVMe drives are expected to be particularly beneficial for I/O intensive applications in datacenter installations. In this paper we identify and analyze the different factors leading to the better performance of NVMe SSDs. Then, using databases as the prominent use-case, we show how these would translate into real-world benefits. We evaluate both a relational database (MySQL) and a NoSQL database (Cassandra) and demonstrate significant performance gains over best-in-class enterprise SATA SSDs: from 3.5x for TPC-C and up to 8.5x for Cassandra. Qiumin Xu, Huzefa Siyamwala, Mrinmoy Ghosh, Manu Awasthi, Tameesh Suri, Zvika Guz, Anahita Shayesteh, Vijay Balakrishnan |
SIGMETRICS | 7 |
| 2015 | Performance analysis of NVMe SSDs and their implication on real world databasesabstractThe storage subsystem has undergone tremendous innovation in order to keep up with the ever-increasing demand for throughput. Non Volatile Memory Express (NVMe) based solid state devices are the latest development in this domain, delivering unprecedented performance in terms of latency and peak bandwidth. NVMe drives are expected to be particularly beneficial for I/O intensive applications, with databases being one of the prominent use-cases. Qiumin Xu, Huzefa Siyamwala, Mrinmoy Ghosh, Tameesh Suri, Manu Awasthi, Zvika Guz, Anahita Shayesteh, Vijay Balakrishnan |
SYSTOR | 7 |
| 2015 | System-Level Characterization of Datacenter ApplicationsabstractIn recent years, a number of benchmark suites have been created for the ``Big Data'' domain, and a number of such applications fit the client-server paradigm. A large volume of recent literature in characterizing ``Big Data'' applications have largely focused on two extremes of the characterization spectrum. On one hand, multiple studies have focused on client-side performance. These involve fine-tuning server-side parameters for an application to get the best client-side performance. On the other extreme, characterization focuses on picking one set of client-side parameters and then reporting the server microarchitectural statistics under those assumptions. While the two ends of the spectrum present interesting results, this paper argues that they are not enough, and in some cases, undesirable, to drive system-wide architectural decisions in datacenter design. Manu Awasthi, Tameesh Suri, Zvika Guz, Anahita Shayesteh, Mrinmoy Ghosh, Vijay Balakrishnan |
ICPE | 4 |
| 2013 | SIMD divergence optimization through intra-warp compactionabstractSIMD execution units in GPUs are increasingly used for high performance and energy efficient acceleration of general purpose applications. However, SIMD control flow divergence effects can result in reduced execution efficiency in a class of GPGPU applications, classified as divergent applications. Improving SIMD efficiency, therefore, has the potential to bring significant performance and energy benefits to a wide range of such data parallel applications. Aniruddha S. Vaidya, Anahita Shayesteh, Dong Hyuk Woo, Roy Saharoy, Mani Azimi |
ISCA | 2 |
| 2006 | Improving the performance and power efficiency of shared helpers in CMPsabstractTechnology scaling trends have forced designers to consider alternatives to deeply pipelining aggressive cores with large amounts of performance accelerating hardware. One alternative is a small, simple core that can be augmented with latency tolerant helpers. As the demands placed on the processor core varies between applications, and even between phases of an application, the benefit seen from any set of helpers will vary tremendously. If there is a single core, these auxiliary structures can be turned on and off dynamically to tune the energy/performance of the machine to the needs of the running application.As more of the processor is broken down into helpers, and additional cores are added to a single chip that can potentially share helpers, the decisions that are made about these structures become increasingly important. In this paper we describe the need for methods that effectively manage these helpers. Our counter-based approach can dynamically turn off three helpers on average while staying within 2% of the performance when running with all helpers. In a multicore environment, our intelligent and exible sharing of helper provides an average 24% speedup compared to static sharing in conjoined cores. Furthermore we show a benefit from constructively sharing helpers among multiple cores running the same application. Anahita Shayesteh, Glenn Reinman, Norman P. Jouppi, Timothy Sherwood, Suleyman Sair |
CASES | 1 |
| 2005 | Reducing the Latency and Area Cost of Core Swapping through Shared Helper EnginesabstractTechnologies scaling trends and the limitations of packaging and cooling have intensified the need for thermally efficient architectures and architecture-level temperature management techniques. To combat these trends, we explore the use of core swapping on microcore architecture, a deeply decoupled processor core with larger structures factored out as helper engines. The microcore architecture presents an ideal platform for core swapping thanks to helper engines that maintain the state of each process in a shared fabric surrounding the cores, reducing the impact of core swapping 43% on average while showing promising thermal reduction. It also has favorable performance when compared to other thermal management techniques. Furthermore, we evaluate alternative approaches to spending the area overhead of the additional microcore, including larger microcores, CMP cores, and SMT cores with different thermal management techniques. Anahita Shayesteh, Eren Kursun, Timothy Sherwood, Suleyman Sair, Glenn Reinman |
ICCD | 1 |
| 2005 | Tornado warning: the perils of selective replay in multithreaded processorsabstractAs future technologies push towards higher clock rates, traditional scheduling techniques that are based on wake-up and select from an instruction window fail to scale due to their circuit complexities. Speculative instruction schedulers can significantly reduce logic on the critical scheduling path, but can suffer from instruction misscheduling that can result in wasted issue opportunities.Misscheduled instructions can spawn other misscheduled instructions, only to be replayed over again and again until correctly scheduled. These "tornadoes" in the speculative scheduler are characterized by extremely low useful scheduling throughput and a high volume of wasted issue opportunities. The impact of tornadoes becomes even more severe when using Simultaneous Multithreading. Misschedulings from one thread can occupy a significant portion of the processor issue bandwidth, effectively starving other threads.In this paper, we propose Zephyr, an architecture that inhibits the formation of tornadoes. Zephyr makes use of existing load latency prediction techniques as well as coarse-grain FIFO queues to buffer instructions before entering scheduling queues. On average, we observe a 23% improvement in IPC performance, 60% reduction in hazards, 41% reduction in occupancy, and 48% reduction in the number of replays compared with a baseline scheduler. Yongxiang Liu, Anahita Shayesteh, Gokhan Memik, Glenn Reinman |
ICS | 2 |
| 2004 | Scaling the issue window with look-ahead latency predictionabstractIn contemporary out-of-order superscalar design, high IPC is mainly achieved by exposing high instruction level parallelism (ILP). Scaling issue window size can certainly provide more ILP; however, future processor scaling demands threaten to limit the size of the issue window.In this study, we propose a dynamic instruction sorting mechanism that provides more ILP without increasing the size of the issue window. In our approach, early in the pipeline, we predict how long an instruction needs to wait before it can be issued, i.e. the waiting time for its operands to be produced. Using this knowledge, the instructions are placed into a sorting structure, which allows instructions with shorter waiting times enter the issue window ahead of those instructions with longer waiting times, preventing long-waiting instructions from clogging the issue queue.The accuracy in predicting instruction waiting times directly determines the effectiveness of our sorting mechanism. While most instructions have deterministic execution latencies, predicting load execution times is more difficult due to cache misses and in-flight loads. Loads are particularly challenging since their execution time can vary significantly. In this study, we examine techniques to predict load execution time accurately, based on data reference history. Yongxiang Liu, Anahita Shayesteh, Gokhan Memik, Glenn Reinman |
ICS | 2 |