VLDB 2026 Research / reviewers in the wild / expert
Mahesh Ketkar
dblp:80/6049 · also Mahesh C. Ketkar
· DBLP profile ↗
15ranked-venue papers
3as first author
8since 2021 · last 2025
0000-0002-8884-5010ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 12 · 3 first-author · 5 since 2021Software engineering, systems software and programming languages · 4 · 4 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | SPIRE: Inferring Hardware Bottlenecks from Performance Counter DataabstractThe persistent demand for greater computing efficiency, coupled with diminishing returns from semiconductor scaling, has led to increased microarchitecture complexity and diversity. Thus, it has become increasingly difficult for application developers and hardware architects to accurately identify low-level performance bottlenecks. Abstract performance models, such as roofline models, help but strip away important microarchitectural details. In contrast, analyses based on hardware performance counters preserve detail but are challenging to implement. This work proposes SPIRE, a novel performance model that combines the accessibility and generality of roofline models with the microarchitectural detail of performance counters. SPIRE (Statistical Piecewise Linear Roofline Ensemble) uses a collection of roofline models to estimate a processor's maximum throughput, based on data from its performance counters. Training this ensemble simply requires sampling data from a processor's performance counters. After training a SPIRE model on 23 workloads running on a CPU, we evaluated it with 4 new workloads and compared our findings against a commercial performance analysis tool. We found that our SPIRE analysis accurately identified many of the same bottlenecks while requiring minimal deployment effort. Nicholas Wendt, Mahesh Ketkar, Valeria Bertacco |
DATE | 2 |
| 2025 | Understanding and Profiling CXL.mem Using PathFinderabstractCXL.mem and the resulting memory pool are promising and gaining great attention. Unlike local memory, CXL DIMMs stay at the I/O subsystem, whose inferior performance can easily impact the processor pipeline and memory subsystem, yielding performance interference, hardware contention, obscure behaviors, and underutilized communication and computing resources. However, our community lacks a tool to understand and profile the CXL.mem protocol execution end-to-end between CPU and remote DIMM. Zerui Guo, Yuebin Bai, Mahesh Ketkar, Hugh Wilkinson, Ming Liu 0027 |
SIGCOMM | 4 |
| 2024 | Aiding Microprocessor Performance Validation with Machine LearningabstractMicroprocessor validation is a complex task that consumes substantial engineering time. Degradation of the system performance that does not affect its functional correctness, is particularly difficult to address given the lack of a golden reference for performance. This work introduces an automated methodology based on machine learning to assist in localizing performance faults, aiming to speed up the validation process. Our results show that, for the injected performance issues, whose average IPC impact is$> 1{\%}$, our technique is able to help localize the exact microarchitectural unit where the degradation occurs$\sim$75% of the time while achieving a top-3 unit accuracy (out of 11 possible locations) of$> 97{\%}$. The proposed setup requires a few seconds to perform a localization inference, leading to a reduced validation time. Erick Carvajal Barboza, Mahesh Ketkar, Paul Gratz, Jiang Hu 0001 |
ISPASS | 2 |
| 2023 | Ditto: End-to-End Application Cloning for Networked Cloud ServicesabstractThe lack of representative, publicly-available cloud services has been a recurring problem in the architecture and systems communities. While open-source benchmarks exist, they do not capture the full complexity of cloud services. Application cloning is a promising way to address this, however, prior work is limited to CPU-/cache-centric, single-node services, operating at user level. Mingyu Liang, Yu Gan 0002, Abhishek Dhanotia, Mahesh Ketkar, Christina Delimitrou |
ASPLOS (2) | 6 |
| 2023 | MQL: ML-Assisted Queuing Latency Analysis for Data Center NetworksabstractData center network (DCN) performance analysis is becoming increasingly critical due to the growing data center scale and proliferation of latency-critical applications. Packetlevel simulators, the de-facto performance evaluation tools, allow accurate modeling of the network and protocols, but they are extremely slow. Simulation of large-scale DCNs with thousands of nodes can take days, making meaningful design space exploration impractical. Analytical techniques, such as queuing theory, can mitigate the scalability problem and offer high accuracy when specific workload assumptions are satisfied. However, their accuracy may decline as these assumptions break, and execution times explode unless designed carefully. To address these challenges, we propose a novel and scalable performance analysis methodology that combines two powerful techniques. First, it uses queuing theory and the maximum entropy (ME) principle to approximate the waiting time in each queue in a DCN. It then finds the end-to-end latency of each flow using traffic input, routing algorithm, and network parameters. This ME-based queuing model can approximate the latency under generalized exponential input traffic and general service distributions. Since its accuracy can degrade as traffic diverges from input and service time assumptions, the second step of the proposed methodology learns and corrects the systematic errors using a regression tree. The resulting ML-assisted technique achieves less than 3% modeling error on average compared to ns-3 simulations. Moreover, the speedup over ns-3 ranges from 100× to 9000× on DCNs with 128 to 1024 nodes. Shruti Yadav Narayana, Jie Tong, Anish Krishnakumar, Nuriye Yildirim, Emily Shriver, Mahesh Ketkar, Ümit Y. Ogras |
ISPASS | 6 |
| 2022 | Mining Patterns From Concurrent Execution TracesabstractThis article proposes a specification mining framework,FlowMiner, that automatically mines patterns from highly concurrent communication traces for system-on-chip (SoC) designs. It addresses the problem of the lack of comprehensive, accurate, and up-to-date specifications necessary to perform rigorous and thorough validation of complex SoC designs. The extracted patterns characterize how components of an SoC design communicate and coordinate with each other to realize various system functions. InFlowMiner, a set of inference rules and optimization techniques are presented to reduce mining complexity. Evaluation of this framework in several experiments shows promising results. Md Rubel Ahmed, Hao Zheng 0001, Parijat Mukherjee, Mahesh Ketkar, Jin Yang 0006 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2021 | Automatic Microprocessor Performance Bug DetectionabstractProcessor design validation and debug is a difficult and complex task, which consumes the lion's share of the design process. Design bugs that affect processor performance rather than its functionality are especially difficult to catch, particularly in new microarchitectures. This is because, unlike functional bugs, the correct processor performance of new microarchitectures on complex, long-running benchmarks is typically not deterministically known. Thus, when performance benchmarking new microarchitectures, performance teams may assume that the design is correct when the performance of the new microarchitecture exceeds that of the previous generation, despite significant performance regressions existing in the design. In this work we present a two-stage, machine learning-based methodology that is able to detect the existence of performance bugs in microprocessors. Our results show that our best technique detects 91.5% of microprocessor core performance bugs whose average IPC impact across the studied applications is greater than 1% versus a bug-free design with zero false positives. When evaluated on memory system bugs, our technique achieves 100% detection with zero false positives. Moreover, the detection is automatic, requiring very little performance engineer time. Erick Carvajal Barboza, Sara Jacob, Mahesh Ketkar, Michael Kishinevsky, Paul Gratz, Jiang Hu 0001 |
HPCA | 3 |
| 2021 | Model Synthesis for Communication Traces of System DesignsabstractConcise and abstract models of system-level behaviors are invaluable in design analysis, testing, and validation. In this paper, we consider the problem of inferring models from communication traces of system-on-chip (SoC) designs. The traces capture communications among different blocks of a system design in terms of messages exchanged. The extracted models characterize the system-level communication protocols governing how blocks exchange messages, and coordinate with each other to realize various system functions. In this paper, the above problem is formulated as a constraint satisfaction problem, which is then fed to a satisfiability modulo theories (SMT) solver. The solutions returned by the SMT solver are used to extract the models that accept the input traces. In the experiments, we demonstrate the proposed approach with traces collected from a transaction-level simulation model of a multicore SoC design and a trace of a more detailed multicore SoC modeled in GEM5. Hao Zheng 0001, Md Rubel Ahmed, Parijat Mukherjee, Mahesh Ketkar, Jin Yang 0006 |
ICCD | 4 |
| 2009 | A microarchitecture-based framework for pre- and post-silicon power delivery analysisabstractVariations in power supply voltage, which is a function of the power delivery network and dynamic current consumption, can affect circuit reliability. Much work has been done to understand power delivery robustness during both the design phase as well as the post-silicon validation phase. Methods applicable at the design phase typically synthesize worst-case current waveforms based on simple current constraints but fail to provide corresponding instruction streams due to their ignorance of the functional aspects of the machine and hence cannot be validated. Approaches used for post-silicon validation are not useful during design, and either rely heavily on available test content which can come from power, performance, or defect testing, and hence are limited in validation potential or employ manually-crafted tests aimed at power delivery, and hence are highly labor-intensive. In this paper, we provide a novel approach to construct processor current waveforms to induce significant droops while at the same time producing instruction streams to achieve those waveforms. We solve the pre-silicon current stimulus generation problem as an optimization problem. The modular framework in this paper utilizes microarchitectural information, current consumption estimates of fine-grained microarchitectural components and a pre-characterized power delivery network to obtain significant droop-inducing current waveforms. The paper further discusses techniques to convert operations associated with these generated waveforms to functional instruction streams. Silicon measurements of such tests run on an industrial microprocessor validate the approach. Mahesh Ketkar, Eli Chiprout |
MICRO | 1 |
| 2009 | Gate Sizing for Cell-Library-Based DesignsabstractWith increasing time-to-market pressure and shortening semiconductor product cycles, more and more chips are being designed with library-based methodologies. In spite of this shift, the problem of discrete gate sizing has received significantly less attention than its continuous counterpart. On the other hand, cell sizes of many realistic libraries are sparse, for example, geometrically spaced, which makes the nearest rounding approach inapplicable as large timing violations may be introduced. Therefore, it is highly desirable to design an effective algorithm to handle this discrete gate-sizing problem. Such an algorithm is proposed in this paper. The algorithm is a continuous-solution-guided dynamic-programming-like approach. A set of novel techniques, such as locality-sensitive-hashing-based solution pruning, is also proposed to accelerate the algorithm. Our experimental results demonstrate that 1) the nearest rounding approach often leads to large timing violations and 2) compared to the well-known Coudert's approach, the new algorithm saves up to 21% in area cost while still satisfying the timing constraint. Shiyan Hu 0001, Mahesh Ketkar, Jiang Hu 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2007 | Comparative Analysis of Conventional and Statistical Design TechniquesabstractWe explore the power benefits of changing a microprocessor path histogram through circuit sizing based on statistical timing analysis and optimization (STAO) versus a deterministic timing approach that uses statistical design to establish a global guardband followed by conventional optimization (SDGG). Using an analytical modeling approach, we quantify the differences in total power between the two approaches while maintaining an equivalent performance distribution. For a relative 1σ random WID stage delay variation of 5% and representative microprocessor critical paths, the analysis indicates that the STAO approach enables ~2% power reduction over the SDGG approach. To achieve a 4% and 6% power reduction through the STAO approach, the process variation needs to increase by a factor of 2x and 4x, respectively. Steven M. Burns, Mahesh Ketkar, Noel Menezes, Keith A. Bowman, James W. Tschanz, Vivek De |
DAC | 2 |
| 2007 | Gate Sizing For Cell Library-Based DesignsabstractAbstract—With increasing time-to-market pressure and short-ening semiconductor product cycles, more and more chips are being designed with library-based methodologies. In spite of this shift, the problem of discrete gate sizing has received significantly less attention than its continuous counterpart. On the other hand, cell sizes of many realistic libraries are sparse, for example, geo-metrically spaced, which makes the nearest rounding approach inapplicable as large timing violations may be introduced. There-fore, it is highly desirable to design an effective algorithm to handle this discrete gate-sizing problem. Such an algorithm is pro-posed in this paper. The algorithm is a continuous-solution-guided dynamic-programming-like approach. A set of novel techniques, such as locality-sensitive-hashing-based solution pruning, is also proposed to accelerate the algorithm. Our experimental results demonstrate that 1) the nearest rounding approach often leads to large timing violations and 2) compared to the well-known Coudert’s approach, the new algorithm saves up to 21 % in area cost while still satisfying the timing constraint. Index Terms—Discretization, dynamic programming (DP), gate sizing, pruning, sparse cell library. I. Shiyan Hu 0001, Mahesh Ketkar, Jiang Hu 0001 |
DAC | 2 |
| 2002 | Standby power optimization via transistor sizing and dual threshold voltage assignmentabstractThis paper presents a novel enumerative approach, with provable and efficient pruning techniques, for dual threshold voltage (Vt) assignment at the transistor level. Since the use of low Vt may entail a substantial increase in leakage power, we formulate the problem as one of combined optimization for leakage-delay tradeoffs under Vt optimization and sizing. Based on an analysis of the effects of these two transforms on the delay and leakage, we justify a two-step procedure for performing this optimization. Results are presented on the ISCAS85 benchmark suite favorably comparing our approach with an existing sensitivity-based optimizer. Mahesh Ketkar, Sachin S. Sapatnekar |
ICCAD | 1 |
| 2000 | Convex delay models for transistor sizingabstractThis paper derives a methodology for developing accurate convex delay models to be used for transistor sizing. A new rich class of convex functions to model gate delay is presented and the circuit delay under such a model is shown to be equivalent to a convex function. The richness of these functions is exploited to accurately model gate delay for modern designs. The delay model is incorporated into a transistor sizing algorithm based on TILOS. The models were characterized by using a set of grid points and then validated using a disjoint data set. The models were found to be within about 10% of SPICE for nearly all of the gate types considered. Also presented are the experimental results of sizing various test circuits. Mahesh Ketkar, Kishore Kasamsetty, Sachin S. Sapatnekar |
DAC | 1 |
| 2000 | A new class of convex functions for delay modeling and itsapplication to the transistor sizing problem [CMOS gates]abstractThis paper derives a methodology for developing accurate convex delay models to be used for transistor sizing. A new rich class of convex functions to model gate delay is presented and the circuit delay under such a model is shown to be equivalent to a convex function. The richness of these functions is exploited to accurately model gate delay for modern designs. Since the delay under this model is a convex function, optimal sizing algorithms based on convex programming techniques are applied with the new delay model. Experimental results demonstrating the accuracy of proposed model are presented along with results of sizing various test circuits. Kishore Kasamsetty, Mahesh Ketkar, Sachin S. Sapatnekar |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |