VLDB 2026 Research / reviewers in the wild / expert
Sanyam Mehta
dblp:141/0166
· DBLP profile ↗
16ranked-venue papers
11as first author
8since 2021 · last 2025
0009-0005-5319-689XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 11 · 7 first-author · 4 since 2021Software engineering, systems software and programming languages · 6 · 4 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | EVeREST-C: An Effective and Versatile Runtime Energy Saving Tool for CPUsabstractPower and energy efficiency are increasingly important challenges within HPC.However, it is still important to achieve these goals while maintaining desired/high application performance.Balancing these goals involves the challenge of precise application characterization.For successful user adoption, this must avoid modifying the application and/or extraneous application profiling, and also be portable to different processors across processor generations and vendors.We propose EVeREST-C to solve these challenges.Everest targets the finer-grained individual application functions for exploiting power/energy saving opportunities via Dynamic Voltage Frequency Scaling (DVFS) in both the core and the uncore, without application-specific knowledge.Since Everest relies on a single standard and accurate performance event, IPS (instructions per second), for its characterization rather than on the (many) performance counters that can differ across platforms, it is portable across processors.Finally, the fine-grained approach enables Everest to additionally save power/energy for select communication (MPI) phases, where appropriate phases are chosen based on both their length and position in the application with regards to the memory/compute boundedness of surrounding user routines.We evaluate Everest using SPEC CPU 2017 and various MPI applications, on Intel and AMD platforms.We find that Everest saves on average 11% more energy for SPEC compared to the baseline and 8% more energy on MPI applications compared to a state-of-the-art solution. Anna Yue, Pen-Chung Yew, Sanyam Mehta |
ICS | 3 |
| 2025 | EVeREST: An Effective and Versatile Runtime Energy Saving Tool for GPUsabstractAmid conflicting demands for ever-improving performance and maximizing energy savings, it is important to have a tool that automatically identifies opportunities to save power/energy at runtime without compromising performance. GPUs in particular present challenges due to (1) reduced savings available from memory bound applications, and (2) limited availability of low overhead performance counters. Thus, a successful tool must address these issues while still tackling the challenges of dynamic application characterization, versatility across processors from different vendors, and effectiveness at making the right power-performance tradeoffs for desired energy savings. Anna Yue, Pen-Chung Yew, Sanyam Mehta |
PPoPP | 3 |
| 2024 | Forward to the Past: An Alternative to Hybrid CPU DesignabstractHybrid CPUs are a response from industry to the increasingly difficult problem of power management in desktop-class processors. Hybrid CPUs comprise a big/performance core for regular tasks and a small/efficiency core primarily for background OS tasks. Although helpful in managing power, the introduction of two different kinds of cores in such a design significantly complicates the scheduling of applications across cores. It also introduces other issues such as ISA compatibility, OS compatibility, many processor SKUs/configurations and the difficulty with designing/validating two independent cores. We propose to simplify with the use of a single core that is application-aware and adjusts the core and the caches both based on observed IPC within application regions and priority across applications (regular versus background tasks). We show that this simpler design is similarly effective in managing system power and saving energy as the existing hybrid CPU s with two different kinds of cores. Sanyam Mehta, Anna Yue |
ISPASS | 1 |
| 2023 | Speculative Register ReclamationabstractLarge number of in-flight instructions were envisioned two decades ago. They are finally happening now. While more in-flight instructions enable higher ILP and therefore better single-thread performance, it comes at a price. The price is larger structures within the core such as the physical register file. In this work, we propose to reduce register file size while maintaining (or even increasing) number of in-flight instructions. We leverage the insight that within loops, where most time is spent in general, most logical registers are redefined in the same or immediate next iteration. The physical registers allocated to most of these logical registers can thus be aggressively and speculatively released at redefinition instead of being released when the redefining instruction finally commits as in current designs. For correct mis-speculation recovery, only registers that are actually used (i.e used without prior redefinition) across iterations need to remain allocated beyond redefinition, leading to much reduced register file pressure. We show that using our design, register file sizes can be reduced by 50% while still achieving a 1.05x performance improvement over existing designs on a variety of applications even when other core resources are kept the same. The power consumption among various core structures in reduced by 26% on average. In addition, the performance improvement jumps to 1.14x when this reduction in register file size is complemented with an increase of other structures within the core. Sanyam Mehta |
HPCA | 1 |
| 2023 | An Application-Oriented Approach to Designing Hybrid CPU ArchitecturesabstractHybrid CPUs have recently launched in desktop and laptop devices with the goal of increasing core count at manageable power consumption. These CPUs contain ‘performance’ and ‘efficiency’ cores, with a dedicated scheduler to assign tasks to cores. We find that these cores are not well-suited for any specific class of applications, and the efficiency cores are often smaller versions of the performance cores. This reduces the efficacy of scheduling tasks to cores. We show that instructions per cycle (IPC) per core serves as a natural metric to divide applications into two distinct classes. We use this division to propose a ‘mountain’ core with large core structures and a lower dispatch/retire width and decreased execution units (and thus lower number of ports) for low IPC applications, and a ‘plateau’ core with small core structures and increased dispatch width and execution units for high IPC applications. These changes tailor the hybrid CPU to the applications being run, allowing us to improve in both power and performance. We find that when compared to a Goldencovelike performance core, the plateau core on average improves performance by 8%, performance per Watt by 14%, and ED2P by 25% for high IPC applications; the mountain core on average improves power by 30%, performance per Watt by 34%, and ED2P by 17% for low IPC applications. Anna Yue, Sanyam Mehta |
ISPASS | 2 |
| 2022 | Software pre-execution for irregular memory accesses in the HBM eraabstractThe introduction of High Bandwidth Memory (HBM) necessitates the use of intelligent software prefetching in irregular applications to utilize the surplus bandwidth. In this work, we propose Software Pre-execution (SPE), a technique that relies on pre-executing a minimal copy of the loop of concern (we call the pre-execution loop) for the purpose of prefetching irregular accesses. This is complemented by the compiler's enforcing a certain prefetch distance through apriori strip-mining of the original loop such that the execution of the pre-execution loop is interspersed with the main loop to ensure timeliness of prefetches. We find that this approach provides natural advantages over prior art such as preservation of loop vectorization, handling short loops, avoiding performance bottlenecks, amenability to threading and most importantly, effective coverage. We demonstrate these advantages using a variety of benchmarks on Fujitsu's A64FX processor with HBM2 memory - we outperform prior art by 1.3x and 1.2x when using small and huge pages, respectively. Simulations further show that our approach holds stronger promise on upcoming processors with HBM2e. Sanyam Mehta, Gary Elsesser, Terry Greyzck |
CC | 1 |
| 2022 | Performance Analysis and Optimization with Little's LawabstractPerformance tools are the bridge between processor architecture and a user. However, with the increasingly complex processor architectures, it is becoming increasingly difficult for the users to comprehend the information generated by the performance tools to help diagnose and fix the performance bottlenecks. In addition, the performance tools are themselves limited in many cases. Finally, there is wide variability in the kind of performance counters provided by the different processor vendors, making performance tools unportable across emerging architectures. In this work, we propose to solve these problems by accurately computing a portable and easily comprehensible performance metric - the (Memory-Level Parallelism) MLP of an application. The observed MLP when seen as a fraction of peak theoretical MLP supported by the host processor provides important guidance on the applicability of various popular program optimizations. Six case studies on three different processors each with a different memory technology show that our metric is both effective in program analysis and provides useful guidance on program optimization. Sanyam Mehta |
ISPASS | 1 |
| 2021 | Variable-Sized Blocks for Locality-Aware SpMVabstractBlocking is an important optimization option available to mitigate the data movement overhead and improve the temporal locality in SpMV, a sparse BLAS kernel with irregular memory reference pattern. In this work, we propose an analytical model to determine the effective block size for highly irregular sparse matrices by factoring the distribution of non-zeros in the sparse dataset. As a result, the blocks generated by our scheme are variable-sized as opposed to constant-sized in most existing SpMV algorithms. We demonstrate our blocking scheme using Compressed Vector Blocks (CVB), a new column-based blocked data format, on Intel Xeon Skylake-X multicore processor. We evaluated the performance of CVB-based SpMV with variable-sized blocks using extensive set of matrices from Stanford Network Analysis Platform (SNAP). Our evaluation shows a speedup of up to 2.62X (with an average of 1.73X) and 2.02X (with an average of 1.18X) over the highly vendor tuned SpMV implementation in Intel's Math Kernel Library (MKL) on single and multiple Intel Xeon cores respectively. Naveen Namashivavam, Sanyam Mehta, Pen-Chung Yew |
CGO | 2 |
| 2016 | WearCore: A Core for Wearable WorkloadsabstractLately, the industry has recognized immense potential in wearables (particularly, smartwatches) being an attractive alternative/supplement to the smartphone. To this end, there has been recent activity in making the smartwatch `self-sufficient' i.e. using it to make/receive calls, etc. independently of the phone. This marked shift in the way wearables will be used in future calls for changes in the core micro-architecture of smartwatch processors. Sanyam Mehta, Josep Torrellas |
PACT | 1 |
| 2016 | TurboTiling: Leveraging Prefetching to Boost Performance of Tiled CodesabstractLoop tiling or blocking improves temporal locality by dividing the problem domain into tiles and then repeatedly accessing the data within a tile. While this reduces reuse, it also leads to an often ignored side-effect: breaking the streaming data access pattern. As a result, tiled codes are unable to exploit the sophisticated hardware prefetchers in present-day processors to extract extra performance. Sanyam Mehta, Rajat Garg, Nishad Trivedi, Pen-Chung Yew |
ICS | 1 |
| 2016 | Variable LiberalizationabstractIn the wake of the current trend of increasing the number of cores on a chip, compiler optimizations for improving the memory performance have assumed increased importance. Loop fusion is one such key optimization that can alleviate memory and bandwidth wall and thus improve parallel performance. However, we find that loop fusion in interesting memory-intensive applications is prevented by the existence of dependences between temporary variables that appear in different loop nests. Furthermore, known techniques of allowing useful transformations in the presence of temporary variables, such as privatization and expansion, prove insufficient in such cases. In this work, we introduce variable liberalization , a technique that selectively removes dependences on temporary variables in different loop nests to achieve loop fusion while preserving the semantical correctness of the optimized program. This removal of extra-stringent dependences effectively amounts to variable expansion, thus achieving the benefit of an increased degree of freedom for program transformation but without an actual expansion. Hence, there is no corresponding increase in the memory footprint incurred. We implement liberalization in the Pluto polyhedral compiler and evaluate its performance on nine hot regions in five real applications. Results demonstrate parallel performance improvement of 1.92 × over the Intel compiler, averaged over the nine hot regions, and an overall improvement of as much as 2.17 × for an entire application, on an eight-core Intel Xeon processor. Sanyam Mehta, Pen-Chung Yew |
ACM Trans. Archit. Code Optim. | 1 |
| 2015 | Improving compiler scalability: optimizing large programs at small priceabstractCompiler scalability is a well known problem: reasoning about the application of useful optimizations over large program scopes consumes too much time and memory during compilation. This problem is exacerbated in polyhedral compilers that use powerful yet costly integer programming algorithms to compose loop optimizations. As a result, the benefits that a polyhedral compiler has to offer to programs such as real scientific applications that contain sequences of loop nests, remain impractical for the common users. In this work, we address this scalability problem in polyhedral compilers. We identify three causes of unscalability, each of which stems from large number of statements and dependences in the program scope. We propose a one-shot solution to the problem by reducing the effective number of statements and dependences as seen by the compiler. We achieve this by representing a sequence of statements in a program by a single super-statement. This set of super-statements exposes the minimum sufficient constraints to the Integer Linear Programming (ILP) solver for finding correct optimizations. We implement our approach in the PLuTo polyhedral compiler and find that it condenses the program statements and program dependences by factors of 4.7x and 6.4x, respectively, averaged over 9 hot regions (ranging from 48 to 121 statements) in 5 real applications. As a result, the improvements in time and memory requirement for compilation are 268x and 20x, respectively, over the latest version of the PLuTo compiler. The final compile times are comparable to the Intel compiler while the performance is 1.92x better on average due to the latter’s conservative approach to loop optimization. Sanyam Mehta, Pen-Chung Yew |
PLDI | 1 |
| 2014 | Multi-stage coordinated prefetching for present-day processorsabstractData prefetching is an important technique for hiding memory latency. Latest microarchitectures provide support for both hardware and software prefetching. However, the architectural features supporting either are different. In addition, these features can vary from one architecture to another. As a result, the choice of the right prefetching strategy is non-trivial for both the programmers and compiler-writers. Sanyam Mehta, Zhenman Fang, Antonia Zhai, Pen-Chung Yew |
ICS | 1 |
| 2014 | Revisiting loop fusion in the polyhedral frameworkabstractLoop fusion is an important compiler optimization for improving memory hierarchy performance through enabling data reuse. Traditional compilers have approached loop fusion in a manner decoupled from other high-level loop optimizations, missing several interesting solutions. Recently, the polyhedral compiler framework with its ability to compose complex transformations, has proved to be promising in performing loop optimizations for small programs. However, our experiments with large programs using state-of-the-art polyhedral compiler frameworks reveal suboptimal fusion partitions in the transformed code. We trace the reason for this to be lack of an effective cost model to choose a good fusion partitioning among the possible choices, which increase exponentially with the number of program statements. In this paper, we propose a fusion algorithm to choose good fusion partitions with two objective functions - achieving good data reuse and preserving parallelism inherent in the source code. These objectives, although targeted by previous work in traditional compilers, pose new challenges within the polyhedral compiler framework and have thus not been addressed. In our algorithm, we propose several heuristics that work effectively within the polyhedral compiler framework and allow us to achieve the proposed objectives. Experimental results show that our fusion algorithm achieves performance comparable to the existing polyhedral compilers for small kernel programs, and significantly outperforms them for large benchmark programs such as those in the SPEC benchmark suite. Sanyam Mehta, Pei-Hung Lin, Pen-Chung Yew |
PPoPP | 1 |
| 2014 | Measuring Microarchitectural Details of Multi- and Many-Core Memory Systems through MicrobenchmarkingabstractAs multicore and many-core architectures evolve, their memory systems are becoming increasingly more complex. To bridge the latency and bandwidth gap between the processor and memory, they often use a mix of multilevel private/shared caches that are either blocking or nonblocking and are connected by high-speed network-on-chip. Moreover, they also incorporate hardware and software prefetching and simultaneous multithreading (SMT) to hide memory latency. On such multi- and many-core systems, to incorporate various memory optimization schemes using compiler optimizations and performance tuning techniques, it is crucial to have microarchitectural details of the target memory system. Unfortunately, such details are often unavailable from vendors, especially for newly released processors. In this article, we propose a novel microbenchmarking methodology based on short elapsed-time events (SETEs) to obtain comprehensive memory microarchitectural details in multi- and many-core processors. This approach requires detailed analysis of potential interfering factors that could affect the intended behavior of such memory systems. We lay out effective guidelines to control and mitigate those interfering factors. Taking the impact of SMT into consideration, our proposed methodology not only can measure traditional cache/memory latency and off-chip bandwidth but also can uncover the details of software and hardware prefetching units not attempted in previous studies. Using the newly released Intel Xeon Phi many-core processor (with in-order cores) as an example, we show how we can use a set of microbenchmarks to determine various microarchitectural features of its memory system (many are undocumented from vendors). To demonstrate the portability and validate the correctness of such a methodology, we use the well-documented Intel Sandy Bridge multicore processor (with out-of-order cores) as another example, where most data are available and can be validated. Moreover, to illustrate the usefulness of the measured data, we do a multistage coordinated data prefetching case study on both Xeon Phi and Sandy Bridge and show that by using the measured data, we can achieve 1.3X and 1.08X performance speedup, respectively, compared to the state-of-the-art Intel ICC compiler. We believe that these measurements also provide useful insights into memory optimization, analysis, and modeling of such multicore and many-core architectures. Zhenman Fang, Sanyam Mehta, Pen-Chung Yew, Antonia Zhai, James B. S. G. Greensky, Gautham Beeraka, Binyu Zang |
ACM Trans. Archit. Code Optim. | 2 |
| 2013 | Tile size selection revisitedabstractLoop tiling is a widely used loop transformation to enhance data locality and allow data reuse. In the tiled code, however, tiles of different sizes can lead to significant variation in performance. Thus, selection of an optimal tile size is critical to performance of tiled codes. In the past, tile size selection has been attempted using both static analytical and dynamic empirical (auto-tuning) models. Past work using static models assumed a direct-mapped cache for the purpose of analysis and thus proved to be less robust. On the other hand, the auto-tuning models involve an exhaustive search in a large space of tiled codes. In this article, we propose a new analytical model for tile size selection that leverages the high set associativity in modern caches to minimize conflict misses. Our tile size selection model targets data reuse in multiple levels of cache. In addition, it considers the interaction of tiling with the SIMD unit in modern processors in estimating the optimal tile size. We find that these factors, not considered in previous models, are critical in developing a robust model for tile size selection. We implement our tile size selection model in a polyhedral compiler and test it on 12 benchmark kernels using two different problem sizes. Our model outperforms the previous analytical models that are based on reusing data in a single level of cache and achieves an average performance improvement of 9.7% and 20.4%, respectively, over the best square (cubic) tiles for the two problem sizes. In addition, the tile size chosen by our tile size selection algorithm is similar to the best performing size obtained through an extensive search, validating the analytical model underlying the algorithm. Sanyam Mehta, Gautham Beeraka, Pen-Chung Yew |
ACM Trans. Archit. Code Optim. | 1 |