VLDB 2026 Research / reviewers in the wild / expert
Manoj Nambiar 0001
dblp:22/5778 · also Manoj K. Nambiar 0001, Manoj Karunakaran Nambiar
· DBLP profile ↗
24ranked-venue papers
0as first author
11since 2021 · last 2026
0000-0002-9001-0629ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 12 · 6 since 2021Systems, architecture and hardware · 4 · 3 since 2021Databases, data management, data science and information retrieval · 3 · 1 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Security and privacy · 1Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LLM4PART: From CallGraphs to Bottleneck Partitions with LLM-Enhanced InsightsabstractModern Python applications often use heterogeneous workloads (mixed compute, data, and I/O) and run on CPUs, GPUs, and accelerators across domains such as language processing, finance, and recommendation systems [1], [2]. Understanding their performance requires analyzing how functions interact and contribute to overall execution. Profiling tools such as cProfile, SnakeViz, and pyinstrument expose function-level bottlenecks, but in practice developers rely on manual inspection or limited profiling, which does not scale well and may not reflect realistic workloads. Moreover, profiling typically reports timings at the function level, making it difficult to understand higher-level execution regions composed of multiple interacting functions. Prior work on automated application decomposition includes tools such as Mono2Micro [3] and ServiceCutter [4], which use runtime traces for modularization, while MLIR and HPVM [1] support backend optimizations after architectural decisions but provide limited support for early performance analysis. In contrast, our approach uses lightweight profiling with cprofile and static call graphs integrated with LLM-based analysis to identify and interpret bottleneck partitions before full-scale deployment. Vyuhita Bonthu, Venkatesh Pasumarti, Ashwin Krishnan, Manoj Nambiar 0001 |
ISPASS | 4 |
| 2025 | CADAEC: Content-Aware Deployment of AI Workloads in Edge-Cloud EcosystemabstractThe rapid growth of edge devices has revolutionized industrial AI applications, including robotics, autonomous systems, and IoT, where real-time processing is essential. These systems face the challenge of managing concurrent, high-volume workloads across resource-constrained edge devices and cloud infrastructure. A major hurdle is optimizing deep learning model deployment across edge-cloud environments in dynamic conditions, particularly where input quality and noise fluctuate under concurrent demands. This paper introduces a novel optimization framework, that addressed these challenges and dynamically selects the most suitable models from a diverse model zoo and determines optimal deployment locations (edge or cloud). The proposed framework leverages content-aware approach to minimize both communication and computation latency while considering hardware limitations and environmental factors. Using a binary linear programming (BILP) approach, our method efficiently balances model distribution of an AI pipeline, maximizing end-to-end performance. We validate this framework on a robotic AI pipeline in real-world, noise-variant environments, comparing content-aware and content-agnostic deployment strategies. Our results demonstrate significant optimization in deployment latency and system performance under high-concurrency conditions, using both content-agnostic and content-aware approaches, highlighting the framework's robustness and scalability. Additionally, we showed the effectiveness of the content-aware approach over the content-agnostic method in optimizing deployment choices and reducing latency, while maintaining the desired qualitative outcomes of the AI pipeline with different communication set up. This makes the content-aware strategy more suitable for complex, real-world environments where input quality and noise vary significantly. Overall, The proposed method presents a compelling solution for optimizing AI pipelines in edge-cloud ecosystems, offering potential for broader applications domains. Ratul Kishore Saha, Sparsh Mittal, Rekha Singhal, Manoj Nambiar 0001 |
ICPE | 4 |
| 2024 | Log Sculptor: Making Logs Great AgainabstractIn production environments where applications generate vast data, minimizing downtime is critical. However, large-scale log data can overwhelm storage and computational resources, making system stability challenging to maintain. Poor logging practices worsen this by creating excessive, irrelevant, or unstructured logs that hinder efficient application-crash detection and resolution. This is an enterprise-wide issue often leading to resource exhaustion and prolonged downtime.To address this, we propose Log Sculptor, a tool that leverages GenAI (Large Language Models) to proactively analyze, identify, and improve logging practices in code. It provides recommendations for log statement adjustments (rectifications, additions, removals) and can apply these to produce code with optimized logs. Log Sculptor operates on a prompt-based approach.We benchmark Log Sculptor on nine open-source codebases annotated by industry practitioners for logging practices. It achieves results comparable to human experts and suggests further refinements to enhance logging. We analyze the cost-benefit trade-offs, demonstrating the potential of GenAI in transforming logging practices, improving debugging efficiency, reducing downtime, and increasing reliability in production environments. Prathit Mehta, Ravi Kumar Singh, Shruti Kunde, Rekha Singhal, Manoj Nambiar 0001 |
IEEE Big Data | 7 |
| 2024 | CAR-LLM: Cloud Accelerator Recommender for Large Language ModelsabstractTransformer-based Large Language Models (LLMs) have garnered significant attention due to their multi-modality and exceptional performance across diverse applications. This surge in popularity has spurred the development of numerous new LLMs and corresponding hardware solutions for efficient deployment. However, deploying LLMs on different accelerators for inference poses a significant challenge due to the vast search space involved. This encompasses considerations such as the number of accelerator chips, instances of accelerators, and the choice of inference framework to meet stringent workload and latency constraints. In this paper, we introduce the Cloud Accelerator Recommender for Large Language Models (CAR-LLM), a framework designed to optimize the deployment of LLMs on available accelerators and hardware across various cloud vendors. CAR-LLM aims to achieve maximum performance with minimal cost by recommending optimal deployment strategies. We outline a cost-effective experimental strategy and investigate key parameters affecting the latency of LLMs on specific hardware. Additionally, we develop a performance model to predict latency and throughput, enhancing deployment efficiency and decision-making for LLM applications. Ashwin Krishnan, Venkatesh Pasumarti, Samarth Inamdar, Arghyajoy Mondal, Manoj Nambiar 0001, Rekha Singhal |
HiPC | 5 |
| 2024 | Scalable and Cost-Effective Edge-Cloud Deployment for Multi-Object Tracking SystemabstractThe proliferation of edge devices has significantly advanced technologies in sectors such as autonomous driving and surveillance. However, deploying machine learning models on these resource-constrained devices presents challenges, including scalability and managing unpredictable workloads, which often hinder real-time performance (both latency and throughput) in edge-only environments. To address these issues, a potential approach is to deploy on an edge-cloud ecosystem, however, it may hinder latency due to communication delay between edge and cloud. We leverage the cloud service Message Queuing Telemetry Transport (MQTT) protocol powered by 5G internet, for communication to manage the latency. We propose a scalable edge-cloud ecosystem specifically designed for a Multi-Camera Sensor-based Multi-Object Tracking (MCS-MOT) pipeline within industrial deployment contexts, ensuring compliance with Service Level Agreements (SLAs). Our approach introduces a novel camera feed content-guided load balancing technique that dynamically manages workloads between edge and cloud. We extract features from incoming camera feed content to inform the load-balancing process efficiently. The load balancer determines the maximum number of concurrent feeds that can be processed at the edge, with the remaining feeds handled by the cloud based on the content of the camera feed. Additionally, we propose a cost model to estimate the expenses of deploying the edge-cloud ecosystem in real-world scenarios with dynamic workloads. Our experimental results show that the proposed SLA-compliant MCS-MOT system significantly outperforms edge-only architectures in terms of latency and throughput for the YOLO-DeepSORT algorithm. We also illustrate the estimated and actual deployment costs of the system, highlighting its cost-effective scalability and optimized real-time performance. Ratul Kishore Saha, Dheeraj Chahal, Rekha Singhal, Manoj Nambiar 0001 |
IC2E | 4 |
| 2022 | A Quantum-Classical Hybrid Method for Image Classification and SegmentationabstractEnormous activity in the Quantum Computing area has resulted in it being considered, together with classical computers, for solving different difficult problems - including those of applied nature. An attempt is made in this work to assemble a pipeline consisting of both quantum and classical processing blocks for the task of image classification and segmentation, keeping in mind the present limitations of the gate-model quantum computers. It is based on the work done for the recent BMW Quantum Computing Challenge related to Automotive Industry. The pipeline handles the real-life sized images as the input and output, rather than the toy-sized examples prevalent in Quantum Computing literature. Apart from breaking down the problem to modules, some of which can be accommodated in the existing simulators and hardware, simplifications of the relevant quantum algorithms are also carried out. Its functionality and utility are brought out by applying it to surface crack segmentation on the popular Kaggle Surface Crack Detection data set. The results of the paper are not only limited to simulations, but also involve running models on Noisy, Intermediate-Scale Quantum processors through AWS. In its entirety, this work may lay the groundwork for quantum/quantum-enhanced image segmentation, and providing interested researchers with a stepping-stone in that direction, as the results demonstrate the efficacy of the proposed method, even with simple versions of the quantum modules. Sayantan Pramanik, M. Girish Chandra, C. V. Sridhar, Aniket Kulkarni, Prabin R. Sahoo, Chethan D. V. Vishwa, Hrishikesh Sharma, Vidyut Navelkar, Sudhakara Poojary, Pranav Shah, Manoj Nambiar 0001 |
SEC | 11 |
| 2022 | High-Performance Deployment of Text Detection Model: Compression and Hardware Platform considerationsabstractNetwork compression is often adopted for high throughput implementation on commercial accelerators. We propose a heuristic based approach to obtain compressed networks with a hardware-friendly architecture as an alternative to conventional NAS algorithms that are computationally expensive. The proposed compressed network introduces 142 $\times$ memory-footprint reduction and provide throughput improvement of 5-8 $\times$ on target hardware platforms, while retaining accuracy within 5% of the baseline trained model. We report performance acceleration on CPU, GPU, and FPGAs for a text detection task. Nupur Sumeet, Karan Rawat, Manoj Nambiar 0001 |
ISPASS | 3 |
| 2022 | Performance Model and Profile Guided Design of a High-Performance Session Based Recommendation EngineabstractSession-based recommendation (SBR) systems are widely used in transactional systems to make personalized recommendations to the end-user. In online retail systems, recommendations-based decisions need to be made at a very high rate especially during peak hours. The required computational workload is very high especially when there is a larger number of products involved. Session Based Recommendation (SBR) models incorporate the learning-based product buying pattern from various user interaction sessions and try to recommend the top-K products, the user is likely to purchase. These models comprise several functional layers that widely vary in their compute and data access patterns. To support high recommendation rates, all these layers need a performance optimal implementation, which can be a challenge given the diverse nature of the computations involved. For this reason, one compute platform - whether it is CPU, GPU, or a Field Programmable Gate Array (FPGA) may not be able to provide an optimal implementation for all the layers. In this paper, we describe performance modeling and profile-based design approach to arrive at an optimal implementation, comprising of the hybrid CPU, GPU, and FPGA platforms for NISER - a session-based recommendation model that avoids popularity bias in recommendations. In addition, the design for the CPU-FPGA hybrid platform is implemented for NISER and we observed that experimental results closely follow the results predicted by the performance model for the implemented deployment option. Ashwin Krishnan, Manoj Nambiar 0001, Nupur Sumeet, Sana Iqbal |
ICPE | 2 |
| 2022 | HLS_Profiler: Non-Intrusive Profiling Tool for HLS based ApplicationsabstractThe High-Level Synthesis (HLS) tools aid in simplified and faster design development without familiarity with Hardware Description Language (HDL) and Register Transfer Logic (RTL) design flow. However, it is not straight forward to associate every line of source code to a clock-cycle of synthesized hardware design. On the other hand, the traditional RTL-based design development flow provides the fine-grained performance profile through waveforms. With the same level of visibility in HLS designs, the designers can identify the performance-bottlenecks and obtain the target performance by iteratively fine-tuning the source code. Although, the HLS development tools provide the low-level waveforms, interpreting them in terms of source code variables is a challenging and tedious task. Addressing this gap, we propose an automated profiler tool, HLS\_Profiler, that provides performance profile of source code in a cycle-accurate manner. The HLS\_Profiler tool is non-intrusive and collectively uses the $łangle$static analysis, dynamic trace$\rangle$ of the source code to present the performance profile report to attribute latent clock cycles to each line of source code. Additionally, we developed a set of associative rules to maintain correctness in performance profile of the HLS design. To verify correctness, we demonstrate the HLS\_Profiler tool on MachSuite Benchmarks and an industry-grade recommendation application. The proposed HLS\_Profiler framework provides visibility into the cycle-by-cycle hardware execution of source-code and aids the designer in making performance-centric decisions. Nupur Sumeet, Deeksha Deeksha, Manoj Nambiar 0001 |
ICPE | 3 |
| 2021 | HLS_PRINT: High Performance Logging Framework on FPGAabstractRecent availability of High-Level Synthesis (HLS) development flow from FPGA vendors like Xilinx [1] and Intel [2] have simplified hardware design development to a great extent. A HLS design development flow includes a compiler which can compile a high-level language, such as C/C++, into the corresponding HDL (Hardware Description Language) representation. Developing applications using HLS would entice enterprise users, given the simplicity of coding in a high-level language. Almost all data center applications running in the data center would log important data. This included the run time contextual data, intermediate steps and final results. This information could serve many purposes like auditing, data for machine learning or just troubleshooting application execution issues. The last requirement is very essential to ensure reduced downtime as availability issues could result in significant loss in business. Logging was not possible with FPGA applications developed even in HLS. Collecting log data necessitates the use of hardware vendor specific modules called integrated logic analyzers (ILA) [3] which require manual integration and support small buffers which limits the amount of data that can be captured.Addressing these issues, we built a logging framework which would enable logging for FPGA implemented applications just as they would for software-based applications. The data would be available to operations staff in similar form as software applications. The logging framework is very similar to the use of printf function available in the C stdio library that comes with standard C-based software compilers.Our logging framework, HLS_PRINT [4], is a hardware-software solution to enable print functionality in the HLS design platforms to bridge the hardware logging gap. We use source-to-source transformations to create HLS synthesizable print constructs. In addition to this, the HLS_PRINT offers a push-button integration into an existing HDL project. The software part of the framework presents logged data in human readable format. Nupur Sumeet, Manoj Nambiar 0001 |
FPL | 2 |
| 2021 | HLS_PRINT: High Performance Logging Framework on FPGAabstractFPGAs have been tipped to be useful for implementing low latency transaction processing systems. Getting computationally powerful over time, they are making their way into enterprise data centers. Another factor is the availability of C compilers for FPGAs as opposed to hardware description languages (HDLs) that requires special skills. However, data center operations staff were concerned about real time troubleshooting in production. Tracing FPGA implemented application execution require special skills and vendor specific tools that can capture limited by the amount of data. To address this, we designed and implemented a logging frame-work on the FPGA. This paper presents the design and implementation of the framework. We present an algorithm that checks and generates alerts for performance overheads introduced due to the use of logging. Finally, experimental results are presented which demonstrate zero or low overhead of the logging framework. Nupur Sumeet, Manoj Nambiar 0001 |
ICPE | 2 |
| 2020 | Recommending in changing timesabstractRecommender systems today face major challenges in keeping up with dynamic customer preferences. Disruptions or sudden changes in the environment affect customer preferences drastically and render historical data ineffective for modeling. With businesses relying heavily on Machine Learning(ML) based recommender systems for catering to customer preferences, the accuracy of timely recommendations gains prime significance. Shruti Kunde, Amey Pandit, Rekha Singhal, Manoj Nambiar 0001, Gautam Shroff |
RecSys | 5 |
| 2019 | Simulation Based Job Scheduling Optimization for Batch WorkloadsabstractWe present a simulation based approach for scheduling jobs that are part of a batch workflow. Our objective is to minimize the makespan, defined as completion time of the last job to leave the system in a batch workflow with dependencies. The existing job schedulers make scheduling decisions based on available cores, memory size, priority or execution time of jobs. This does not guarantee minimum makespan since contention for resources among concurrently running jobs are ignored. In our approach, prior to scheduling batch jobs on physical servers, we simulate the execution of jobs using a discrete event simulator. The simulator considers available cores and available memory bandwidth on distributed systems to accurately simulate the execution of jobs using resource contention models in a concurrent run. We also propose simulation based job scheduling algorithms that use underlying contention models and minimize the makespan by optimally mapping jobs onto the available nodes. Our approach ensures that job dependencies are adhered to during the simulation. We assess the efficacy of our job scheduling algorithms and contention models by performing experiments on a real cluster. Our experimental results show that simulation based approach improves the makespan by 15% to 35% depending on the nature of workload. Dheeraj Chahal, Benny Mathew, Manoj Nambiar 0001 |
ICPE | 3 |
| 2018 | High Performance Distributed In-Memory Architectures for Trade Surveillance SystemabstractWith rapid growth of world economy, the user activity in the capital markets is increased. This results into large number of transactional activities in trading systems. Hence, the trade surveillance system with low latency and high throughput is needed to monitor such a large amount of data in order to improve user experience by reducing discrepancies and frauds. In-memory technology reduces this latency by processing as well as caching data in main memory thereby removing the overhead of disk access. Currently, open-source frameworks such as Apache Ignite, Apache Flink and Kafka Streams provides in-memory streaming and caching functionalities along with scalability and fault-tolerant features. The paper talks about Trade Surveillance System (TSS), which includes Complex Event Processing (CEP). Here we discuss design, implementation and tuning of three different high-performance architectures for trade surveillance system using Ignite, Flink and Kafka Streams as in-memory streaming technologies. Paper also compares system throughput, support for fault tolerance and effect of caching on streaming throughput for all three architectures. Based on experiments, it is seen that Ignite outperforms Flink and Kafka Streams in CEP based streaming. Flink is more reliable considering fault-tolerance and event-time processing at streaming layer compared to Ignite. Though Kafka Streams also provides fault-tolerance and event-time processing out of the box, it shows high latency due to disk based processing. Rishikesh Bansod, Sanket Kadarkar, Rupinder Virk, Mehul Raval, Rushikesh Rashinkar, Manoj Nambiar 0001 |
ISPDC | 6 |
| 2017 | PerfExt++: Performance Extrapolation of IO Intensive WorkloadsabstractWe present a tool, PerfExt++, for cross platform performance extrapolation of IO intensive workloads. The tool is based on trace and replay mechanism and has the capability to record and replay the temporal and spatial characteristics of IO workloads. We show the design and implementation of PerfExt++, which requires minimal intervention from the user to predict and extrapolate performance metrics across platforms. Dheeraj Chahal, Mukund Kumar 0001, Manoj Nambiar 0001 |
ICPE | 3 |
| 2017 | Cloning IO Intensive Workloads Using Synthetic BenchmarkabstractPerformance evaluation of an enterprise application on multiple storage systems of interest called target systems, is a time consuming and costly process. Moreover, it is increasingly challenging to predict the performance at higher concurrencies (no. of users) on target systems when the application is migrated from the low performance source system where the application is currently deployed. Dheeraj Chahal, Manoj Nambiar 0001 |
ICPE | 2 |
| 2017 | Service demand modeling and performance prediction with single-user tests
Ajay Kattepur, Manoj Nambiar 0001 |
Perform. Evaluation | 2 |
| 2016 | Predicting SQL Query Execution Time for Large Data VolumeabstractIn a production system, increase in data size will increase the execution time of the application's SQL queries and degrade its performance. Tuning SQL queries in production requires additional efforts and cost. Time constraints during application development do not permit testing SQL queries with high data volumes. Having the capability to predict SQL query execution time for large data volumes can alert the developers to tune queries or database design upfront, in such scenarios. Application developers may use 'cost' of SQL query as given by optimizer based relational databases to estimate the SQL query execution time for large data sizes. However, the 'cost' based models may lead to large estimation errors as discussed in this paper. We have presented a modular approach of estimating SQL query execution time for high data volumes using measurements at low data volume. A compound SQL query execution plan is mapped to a sequential execution of a set of elementary steps. The execution time of a SQL query in isolation is predicted as summation of estimated execution time of all its elementary steps. We have built analytical models for estimating execution time of different IO access, DB cache access and SQL operators as function of data size for each such step. The proposed model dynamically adapts itself to the structure of the query execution plans and characteristics of the underlying hardware. We have evaluated the model by generating synthetic queries for all combinations of elementary steps for a range of data sizes. The model has also been validated with TPC-H benchmarks and three real life applications. The proposed model shows an average prediction error to be within 10%. Rekha Singhal, Manoj Nambiar 0001 |
IDEAS | 2 |
| 2016 | Performance Extrapolation of IO Intensive Workloads: Work in ProgressabstractPerformance prediction of an application before migrating from a source system and deploying on the target system is a challenging but important task. In this paper, we present a method for predicting the performance of an IO intensive multithreaded enterprise application workload on target systems connected to advanced storage devices. Our approach is an extension of well-known trace and replay method. We extract traces of IO intensive enterprise workloads representing temporal and spatial characteristics (e.g. read and write requests) on the source system where application is currently deployed. These traces are replayed on the system of interest called target system. The experimental results presented demonstrate the effectiveness and accuracy of this method. Dheeraj Chahal, Rupinder Virk, Manoj Nambiar 0001 |
ICPE | 3 |
| 2016 | Maximum Likelihood Estimation of Closed Queueing Network Demands from Queue Length DataabstractResource demand estimation is essential for the application of analyical models, such as queueing networks, to real-world systems. In this paper, we investigate maximum likelihood (ML) estimators for service demands in closed queueing networks with load-independent and load-dependent service times. Stemming from a characterization of necessary conditions for ML estimation, we propose new estimators that infer demands from queue-length measurements, which are inexpensive metrics to collect in real systems. One advantage of focusing on queue-length data compared to response times or utilizations is that confidence intervals can be rigorously derived from the equilibrium distribution of the queueing network model. Our estimators and their confidence intervals are validated against simulation and real system measurements for a multi-tier application. Weikun Wang, Giuliano Casale, Ajay Kattepur, Manoj Nambiar 0001 |
ICPE | 4 |
| 2016 | Recovery From Software Failures Caused by MandelbugsabstractSoftware failures are still a major concern in mission- and enterprise-critical contexts, despite significant efforts spent in software testing. In fact, while software testing is effective against easily-reproducible bugs (Bohrbugs), it is considerably less suitable for dealing with bugs that lead to hard-to-reproduce failures (Mandelbugs). On the positive side, the elusive nature of Mandelbugs provides opportunities for failure recovery, which are investigated in this paper. Based on real cases of Mandelbugs in eleven Information Technology (IT) systems running in production, the paper proposes a model that describes the recovery processes in IT systems. It then presents closed-form expressions, and a numerical analysis, of the mean time to recovery, and the software (un)availability. This analysis allows the designer to compare recovery strategies, as well as to determine the parameters having a high influence on the efficacy of recovery from failures caused by Mandelbugs. Michael Grottke, Dong Seong Kim 0001, Rajesh K. Mansharamani, Manoj Nambiar 0001, Roberto Natella, Kishor S. Trivedi |
IEEE Trans. Reliab. | 4 |
| 2013 | A Tutorial On Modelling Call Centres Using Discrete Event SimulationabstractArriving at an optimal schedule for the staff and determining their required skills in a call centre is imperative to balance the conflicting requirements of delightful customer experience, high employee satisfaction and low cost. Due to the complex nature of modern call centres, simulation modelling is increasingly being used to predict their performance. We have modelled a call centre using our in-house discrete event simulation tool called DESiDE. This paper describes how every component of call centres were modelled as simulation resources. This paper also describes the changes that had to be made to DESiDE in order to handle the special requirements of call centre modelling and also the metrics used by call centres. Benny Mathew, Manoj Nambiar 0001 |
ECMS | 2 |
| 2011 | Recovery from Failures Due to Mandelbugs in IT SystemsabstractSeveral studies have been carried out on software bugs analysis and classification for life and mission critical systems, which include reproducible bugs called Bohrbugs, and hard to reproduce bugs called Mandelbugs. Although software reliability in IT systems has been studied for years, there are only a few formal analytic models for recovery from Mandelbugs. This paper discusses in detail several real cases of Mandelbugs and presents a simple flowchart which describes the recovery processes implemented in IT systems for a large variety of Mandelbugs. The flowchart is based on more than 10 IT systems that are running in production. The paper then presents a closed-form expression of the mean time to recovery from these bugs. Measures of interest including mean time to recovery and system unavailability are computed. A numerical and parametric sensitivity analysis of the model parameters are carried out. This analysis allows the designer to find out important parameter(s) for the recovery from failures due to Mandelbugs. Kishor S. Trivedi, Rajesh K. Mansharamani, Dong Seong Kim 0001, Michael Grottke, Manoj Nambiar 0001 |
PRDC | 5 |
| 2009 | DCPE Rollout: Scaling Performance Engineering Training and Certification across a Very Large EnterpriseabstractPerformance engineering is a badly needed skill for implementing and running IT systems, but performance engineers are hard to find in the market. This paper presents our experiences in rolling out training and certification in a first level course on performance engineering across a large enterprise. We present data and lessons learned on the nominations for the rollout, the design and analysis of theory and practical exams, and the methods used to ensure fairness in a rollout spanning hundreds of nominations. We present results that show trainees to perform exceeding well much against conventional wisdom, and we also show how the success of the rollout has led to a number of beneficial initiatives for the company. Rajesh K. Mansharamani, Arunava Bag, Kishor Gujarathi, Kunal Gupta, Amol Khanapurkar, Manoj Nambiar 0001, Mehul Raval |
CSEE&T | 6 |