Santonu Sarkar

dblp:00/1826 · DBLP profile ↗
← Back
52ranked-venue papers
16as first author
19since 2021 · last 2026
0000-0001-9470-7012ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 21 · 7 first-author · 6 since 2021Systems, architecture and hardware · 18 · 8 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 1 first-author · 1 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-authorComputer networks · 1Databases, data management, data science and information retrieval · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 PLCEQ: Behavioural Equivalence Checking for Industrial PLC Software Migration
Santonu Sarkar, Avijit Mandal, Raoul Praful Jetley
ENASE (1)1
2026 PowerQuant: Architecture-Agnostic GPU Power Estimation via Quantile Regression
abstract
Accurate prediction of NVIDIA GPU power consumption remains challenging due to rapid architectural evolution. Existing machine-learning–based power models are tightly coupled to specific GPU architectures and degrade sharply on unseen platforms, requiring retraining and extensive power measurements, which hinder scalability. This paper presents a quantile-regression–based GPU power prediction framework that enables architecture-agnostic power estimation using static analysis-based compile-time CUDA kernel features. The key insight is that architectural changes primarily induce systematic shifts in power scale, while the relative ordering of kernel power demands remains preserved. By learning power quantiles that capture this ordering and mapping them to new GPUs through one-time calibration, the proposed approach mitigates cross-architecture distribution shift. Extensive evaluation across multiple NVIDIA GPU generations shows that, on unseen architectures, the proposed method improves prediction accuracy by up to 30–50% over existing regression models, while maintaining comparable accuracy in in-distribution settings. The resulting low-overhead, generalizable power estimates make the approach practical for power-aware scheduling, energy budgeting, and sustainability-oriented resource management in large HPC systems.
Aditya Challa, Tanish Desai, Gargi Alavani Prabhu, Snehanshu Saha, Santonu Sarkar
HPDC5
2025 Antarbhukti: Verifying Correctness of PLC Software During System Evolution
Soumyadip Bandyopadhyay, Santonu Sarkar
ATVA2
2025 Predicting Executability and Performance of CNN Kernels on Tenstorrent Hardware Using Machine Learning
abstract
Maximizing the performance of deep learning models on AI accelerators like Tenstorrent Wormhole requires precise control over hardware resources such as compute cores and onchip memory. Operations such as convolution expose a range of tunable parameters, such as parallelization strategies, buffer sizes, compute datatypes, and memory hierarchies (SRAM vs. DRAM) that involve trade-offs between performance, memory usage, and numerical accuracy. Tenstorrent's TT-NN API reimplements PyTorch's conv2d while exposing all these lowlevel controls, allowing fine-grained control, but presenting a steep learning curve for users familiar with PyTorch's high-level abstractions. We propose a predictive software layer that facilitates the execution of conv2d on Tenstorrent hardware by modeling the relationship between input tensors, configuration parameters, and execution outcomes. We train two machine learning models: one predicts execution success with 99.82% accuracy, and the other estimates run-time performance with an$\mathrm{R}^{2}$score of 0.9984. These models enable automated hardware-aware tuning (TuneNet) of the TT-NN configuration space, reducing trial and error, avoiding out-of-memory errors, and improving performance. TuneNetselected configurations reduce conv2d latency by 13.29% compared to TT-NN defaults on Tenstorrent Wormhole N150s, while incurring minimal computational overhead (9s). This work introduces the first hardware-aware autotuning pipeline for Tenstorrent accelerators, significantly reducing tuning overhead while improving performance.
Param Gandhi, Sharvil Potdar, Nayan Gogari, Gargi Alavani Prabhu, Sankar Manoj, Santonu Sarkar
HiPC6
2025 Adaptive GPU Power Capping: Balancing Energy Efficiency, Thermal Control and Performance
abstract
As GPUs become increasingly popular in commodity hardware as well as High Performance Computing(HPC) systems, the need for sustainable computing is more critical. This work addresses the challenge of identifying the optimal operating power for GPUs to minimize energy consumption and operational temperature while incurring only minimal performance overhead. We propose a machine learning-based solution that leverages tree-based models to predict the optimal GPU power cap using key system parameters, including GPU utilization, Memory utilization, Temperature, and Frequency. Our experimental results demonstrate that our model can achieve a maximum energy saving of 12. 87% and a temperature reduction of 11. 38%, with only a 2.69% increase in execution time. These findings highlight the potential of our approach to enhance energy efficiency and thermal management in GPU-based systems, paving the way for more sustainable computing practices.
Tanish Desai, Jainam Shah, Gargi Alavani Prabhu, Snehanshu Saha, Santonu Sarkar
HPDC5
2024 2024 IEEE International Conference on Cloud Computing, Message from the Chairs
abstract
We are delighted to welcome all participants to the 2024 IEEE International Conference on Cloud Computing (CLOUD 2024), which is taking place in the beautiful city of Shenzhen, China, from July 7th to 13th.
Tevfik Kosar, Krishnan Venkateswaran, Shangguang Wang, Seetharami Seelam, Santonu Sarkar, Xuanzhe Liu
CLOUD5
2024 Estimating Power Consumption of GPU Application Using Machine Learning Tool
abstract
As Graphic Processing Units (GPU)s play an increasingly important role in High-Performance Computing (HPC) and data-intensive Machine Learning (ML) tasks, accurate power prediction is essential. Traditional methods, relying on architecture-specific models like DVFS and hardware counters, limit cross-architecture applicability of these models. We propose a static analysis framework that predicts an application's power usage across different NVIDIA GPU architectures without execution. Extensive experiments with state-of-the-art ML approaches show promising results, demonstrating generalizability in predicting power consumption for a newer architecture without the need for complete retraining.11This research is partially supported by the New Faculty Seed Grant of BITS Pilani under Grant No.NFSG/GOA/2023/G0916.
Gargi Alavani Prabhu, Tanish Desai, Sharvil Potdar, Nayan Gogari, Snehanshu Saha, Santonu Sarkar
ICTAI6
2023 Data Flow as Code: Managing Data Flow in an Industrial Hierarchical Edge Network
abstract
For real-time analytics and insight into an industrial production process, it is necessary to deploy a hierarchy of edge nodes that collect massive amounts of telemetry data from many field devices without violating data security constraints in the process core network. Depending on the analytics use case, the requirement for the data at an edge can vary from a hot data stream, warm aggregated data, and cold historical data. Data availability can be severely impacted when communication among edge nodes and devices is interrupted due to network failures. Data backfilling, along with current data transmission with low bandwidth in the process core network is the challenge. This paper proposes a concept for data management called Data Flow as Code (DFaC), which accommodates the requirements of different analytical applications and leverages the advantages of different methods, such as data aggregation, compression, and redundancy removal, and allows configuring edges based on data fingerprints (or data characteristics) to achieve maximum benefit with respect to bandwidth and the quality of the data. Also, we propose DataTrait, which shows a modular implementation of DFaC. We demonstrate the results of extensive numerical experiments conducted with an Azure hierarchical edge setup with open-source Pronto data and Azure predictive maintenance data. Only < 12% of the actual data needs to be backfilled when the connection is restored without compromising data quality, saving huge bandwidth and processing time.
Madapu Amarlingam, AVSK Harish, Santonu Sarkar, Jan-Christoph Schlake 0001
ETFA3
2023 Inspect-GPU: A Software to Evaluate Performance Characteristics of CUDA Kernels Using Microbenchmarks and Regression Models
Gargi Alavani Prabhu, Santonu Sarkar
ICSOFT2
2023 Efficient anomaly identification in temporal and non-temporal industrial data using tree based approaches
Jyotirmoy Sarkar, Snehanshu Saha, Santonu Sarkar
Appl. Intell.3
2022 Modeling Error Propagation in a Modular Plant
abstract
Modular industrial process plants build a production system by integrating a set of predesigned modules supplied by different vendors. The plant process engineers do not have access to the internals of these predesigned modules. Even when a predesigned module is well-tested, the composition of a set of modules can be vulnerable to many unforeseen scenarios, resulting in system-level failures. Debugging the cause of failure becomes extremely difficult since the implementation details are not available to the engineers. This paper proposes a technique by which it models the abstract functionality of each module as a state machine. We model the plant execution as a set of communicating state machines, making it possible to perform error propagation analysis and generate a set of test cases during the engineering phase.
Santonu Sarkar, Nicolai Schoch, Mario Hoernicke
ETFA1
2022 Design of a Validator for Module Type Packages
abstract
Modular plants build a production system by integrating a set of pre-designed modules. Integration of these modules, supplied by different vendors, is performed using various design tools during the engineering phase. The integration process must perform a rigorous validation of the correctness of a module specification (MTP). Otherwise, the integration process can fail without providing enough failure details. Consequently, Such a failure at the later stage can significantly impact implementation, testing, integration, and SAT. In this paper, we describe a validator tool, that allows a plant designer to define a set of invariants that must be satisfied so that an MTP can be deemed fit for integration. We have tested the validator on a set of MTPs and reported our findings. We expect that the use of such a validator can significantly reduce the possibility of introducing errors during the engineering phase.
Santonu Sarkar, Katharina Stark, Mario Hoernicke
IECON1
2022 Performance modeling of graphics processing unit application using static and dynamic analysis
abstract
Summary Graphics processing units (GPUs) have become an integral part of high‐performance computing to achieve an exascale performance. Understanding and estimating GPU performance is crucial for developers to design performance‐driven as well as energy‐efficient applications for a given architecture. This work presents a model developed using a static analysis of CUDA code to predict the execution time of NVIDIA GPU kernels without the need for running it. Here a PTX code is statically analyzed to extract instruction features, control flow, and data dependence. We propose a scheduling algorithm that satisfies resource reservation constraints to schedule these instructions in threads across streaming multiprocessors (SMs). We use dynamic analysis to build a set of memory access penalty models and use these models in conjunction with the scheduling information to estimate the execution time of the code. We present the experimental results which support that this approach works across architectures of NVIDIA GPUs. We first tested our model on two Kepler machines, where the mean percentage error (MPE)/mean absolute percentage error (MAPE) was 8.88%/28.3% for Tesla K20 and 5.66%/29.4% for Quadro K4200. We further tested the model on Maxwell and Pascal architectures and recorded the MPEs/MAPEs to be 10.64%/47.8% and %/28.5%, respectively.
Gargi Alavani Prabhu, Santonu Sarkar
Concurr. Comput. Pract. Exp.2
2022 Architecture of a time-sensitive provisioning system for cloud-native software
abstract
Abstract Application development paradigms and composition of technology services are decisively moving in the direction of hybrid and multi clouds. Enterprises are stitching new cloud‐native business models that leverage containerized multi‐tier microservice architecture, heterogeneity of cloud deployment models, and diversity of cloud providers. Scalability and resiliency are key components of this new world architecture, but these are also functions of the predictability of provisioning the underpinning compute instances on cloud. Thus, a major challenge to surmount before complex multicloud aware applications can be designed is the problem of unpredictable latencies associated with the provisioning of compute services on cloud. In the first part of this article, we develop a technique for time‐sensitive provisioning of virtual compute on demand, while also allowing deprovisioning on demand. Using the technique we propose, a cloud broker will be able to operate on a pool of reserved instances sourced from cloud providers and multiplex them profitably across cloud customers with associated provisioning time guarantees, but without usage commitment restrictions. We articulate this challenge in the form of theReserved Instance Allocation Problem(RIAP), which we first prove to be NP‐hard. We then design a heuristic‐based method to solve the intractable RIAP in polynomial time. We evaluate the effectiveness of our heuristic‐based mechanism through a combination of deep simulations and practical validation on mainstream public clouds. We demonstrate that our algorithm consistently yields a high profit‐to‐investment ratio for a broker who seeks to operate a commerce of virtual machines with time‐sensitive provisioning. In the second part of this article, we tackle the problem of unpredictable provisioning latencies in bare metal commerce. We build and evaluate an allocation model to calculate optimal supporting bare metal inventory to maximize cost‐sensitive fulfillment of bare metal provisioning requests in a time‐sensitive manner.
Sreekrishnan Venkateswaran, Adwait Bauskar, Santonu Sarkar
Softw. Pract. Exp.3
2021 Automatic Control Code Generation from SAMA Specification
abstract
Industrial process automation engineering implements a control code of a process plant manually. This paper proposes an approach that reads a control specification as a multi-page graphical document and implements control code as a Control logic diagram (CLD) for a target controller. We use a novel vector image processing-based approach to extract entities, equipment blocks, and the control flow from this document. We represent the control flow information as language-agnostic intermediate equipment and control flow graph. The translator traverses the graph and applies a set of mapping rules to generate the control logic. A preliminary analysis reveals that our approach has the potential of saving a significant amount of manual effort to generate a CLD from such a specification.
Santonu Sarkar, Chandrika K. R.
ETFA1
2021 Clustering, Separation, and Connection: A Tale of Three Characteristics
abstract
In large and complex software development ecosystems, developers collaborate in multiple dimensions. How characteristics of such collaboration vary over time can offer insights about the dynamics of large scale software development. In this paper we have constructed networks of developers who co-comment on and co-change units of code, and analysed the patterns of variation of clustering, connection, and separation between such developers over time, using development data from a large open source system. Though clustering, connection, and separation essentially represent facets of the same collaboration activities, we found them exhibiting distinct time-varying characteristics. The variation of clustering indicate that developers congregate more closely towards the beginning of the project when the architecture of the system is not yet stable and they need to reach out to one another to fulfil their collective responsibilities. However, separation between developers shows a quick rise and then a saturation around a particular value. Developer connection continues rise throughout the observation period. Using time-series analysis, our results allow us to derive insights on the evolutionary trends in large scale software development and can inform the tuning of tools and process towards effective team assembly and governance.
Subhajit Datta, Aniruddha Mysore, Haziqshah Wira, Santonu Sarkar
ICSME4
2021 d-BTAI: The Dynamic-Binary Tree Based Anomaly Identification Algorithm for Industrial Systems
Jyotirmoy Sarkar, Santonu Sarkar, Snehanshu Saha, Swagatam Das
IEA/AIE (2)2
2021 Identification of Modules from Graphical Control Specification
abstract
Industrial process automation engineering is embracing a modular approach instead of monolithic process plants. In this paper we propose a tool to identify a set of possible modules from a monolithic control specification in the form of a Scientific Apparatus Makers Association (SAMA) document in order to significantly eliminate manual engineering effort. The proposed approach performs vector image processing to extract entities, equipment blocks, and the control flow from the document and store the information as a labeled property graph. Next, we derive a set of possible modules so that the process plant can be implemented as a composition of a set of modules. We perform an unsupervised, K-Means clustering approach to partition the graph where each partition becomes a potential module. Each module consists of a set of equipment blocks, IOLists, and a service description. We leverage a proprietary module design tool to import each cluster as a module skeleton and allow the developer to complete the module description. We have tested the efficacy of this approach on two large SAMA documents and verified the quality of clustering with the domain experts.
Santonu Sarkar
IECON1
2021 Fitness-Aware Containerization Service Leveraging Machine Learning
abstract
Containerized deployment of microservices has gained immense traction across industries. To meet demand, traditional cloud providers offer container-as-a-service, where selection of the container and containerization of workloads remain developer’s responsibility. This task is arduous for a developer since the choice of containers across different cloud providers is many. Furthermore, there does not exist any mechanism using which one can compare and contrast the capabilities of containers across different providers. In this scenario, we envisage the need for a smart cloud broker that can automatically deploy a chosen IT service into the best-fit container environment mapped to performance requirements, from among the set of available underpinning brokered container hosting systems spread across multiple cloud providers. We propose a novel fitness-aware containerization-as-a-service to achieve this. We show why a best-fit container selection process is operationally complex and time consuming, and how we heuristically prune the associated decision tree in two phases so that it becomes viable to implement this as an on-demand service. We propose a new metric called fitness quotient ($FQ$) to evaluate containers obtained from heterogeneous providers. We leverage machine learning techniques to inject automation into these two phases: unsupervised K-Means clustering in the first-level build-time phase to accurately classify IaaS cost and performance data, and polynomial regression during the second-level provisioning-time phase to discover relationships between SaaS performance and container strength. We also show that the utility of the framework that we propose is not limited to the container fitness use case that we analyze in this paper; rather it can be generalized to address a class of problems where overall time and cost complexity for provisioning-time decision making needs to be controlled under a given set of constraints.
Sreekrishnan Venkateswaran, Santonu Sarkar
IEEE Trans. Serv. Comput.2
2020 Analyzing Risky Behavior in Traffic Accidents
abstract
Among all the transportation systems that people use, the public traffic-ways are most common and dangerous resulting in a significant number of fatalities per day worldwide. Statistics have shown that the mortality rates related to traffic accident are more among youth. Although various road safety strategies and rules are developed by the government and law-enforcement agencies to combat the situation, these methods mainly target design, operation, and usability of traffic-ways. Most of the recent data-driven analysis papers model the traffic patterns or predict accidents from the past data. In this paper, we consider a comprehensive, year long fatality analysis reporting system (FARS) data to analyze the role of various factors related to humans, weather and physical conditions (e.g., road surface, light condition etc.) involved in traffic accidents. We build an intelligent risk prediction model that can help decision-makers to ensure road safety. The proposed model estimates (i.) the accident risk over a future time frame, and (ii.) the risk associated with the drivers present on the traffic-way based on the driver's behavior, history, environmental conditions and physical conditions related to traffic-way.
Mayank Chaudhari, Santonu Sarkar, Divyasheel Sharma
SMC2
2020 Analysis, Evaluation, and Assessment for Containerizing an Industry Automation Software
abstract
Container-based virtualization is becoming a preferred choice to deploy services since it is lightweight and supports on-demand scalability as well as availability. The Process Automation Industry has accepted this technology to make their applications service oriented. However, container-based microservice architecture is effective only when the original software strictly followed modularity principles during its design. In this article, we share our learning of converting a distributed software to a microservice-based architecture using containers. Though the existing system has a modular design and deployed as distributed components, analysis of the current architecture shows that the application is monolithic (though modularized) and the components are strongly coupled in an indirect manner. As a result, it turns to be impossible to attain microservice-based architecture without changing the architecture. Next, we propose a microservice-based containerized TO-BE architecture of the application, and demonstrate that this TO-BE architecture does not incur any significant overhead. Finally, we propose a set of recommendations that the practitioners can follow to convert a monolithic application to a containerized architecture.
Santonu Sarkar, P. P. Abdulla, Srini Ramaswamy
SMC1
2020 Assessing Invariant Mining Techniques for Cloud-Based Utility Computing Systems
abstract
Likely system invariants model properties that hold in operating conditions of a computing system. Invariants may be mined offline from training datasets, or inferred during execution. Scientific work has shown that invariants' mining techniques support several activities, including capacity planning and detection of failures, anomalies and violations of Service Level Agreements. However their practical application by operation engineers is still a challenge. We aim to fill this gap through an empirical analysis of three major techniques for mining invariants in cloud-based utility computing systems: clustering, association rules, and decision list. The experiments use independent datasets from real-world systems: a Google cluster, whose traces are publicly available, and a Software-as-a-Service platform used by various companies worldwide. We assess the techniques in two invariants' applications, namely executions characterization and anomaly detection, using the metrics of coverage, recall and precision. A sensitivity analysis is performed. Experimental results allow inferring practical usage implications, showing that relatively few invariants characterize the majority of operating conditions, that precision and recall may drop significantly when trying to achieve a large coverage, and that techniques exhibit similar precision, though the supervised one a higher recall. Finally, we propose a general heuristic for selecting likely invariants from a dataset.
Antonio Pecchia, Stefano Russo 0001, Santonu Sarkar
IEEE Trans. Serv. Comput.3
2019 Time-Sensitive Provisioning of Bare Metal Compute as a Cloud Service
abstract
A modern cloud service that is getting popular is the supply of bare metal servers on demand to consumers who need to run high-performance algorithms for short periods of time. However, provisioning-time latencies are unpredictable in bare metal commerce, primarily because suppliers cannot exploit virtualization levers to adjust and optimize capacity. This impacts bare metal service leverage because, in the face of indeterministic fulfillment times, consumers often pre-provision for peak demand, which goes against the fundamental tenets and advantages of cloud adoption. To address this, we advocate that a cloud service broker offers time-sensitive bare metal provisioning services by building and maintaining an inventory of bare metal servers. This, however, leads to inventory optimization and profit maximization problems that resemble what traditional supply chains face, but with characteristics unique to on-demand compute economics. In this paper, we build a brokered bare metal supply chain model that identifies impacting variables, their characteristics, and inter-relationships. We argue and demonstrate that, given complex inter-variable relationships and environmental uncertainties, simulation runs are needed to complete the model and yield recommendations to optimize inventory and maximize profits.
Sreekrishnan Venkateswaran, Santonu Sarkar
CLOUD2
2019 Improving Safety in Collaborative Robot Tasks
abstract
In recent times, there has been significant interest in collaborative robots where the tasks performed by a robot are non-repetitive and complex, and humans and robots share an overlapping workspace. In such a case, the robot controller must necessarily be safety-aware. In this paper, we propose a mechanism to evaluate a robot program and compute a safety score for each action that the robot is about to perform. To this end, we have implemented a code analyzer that examines the robot's Move instructions and assigns a safety score. A subjective logic based approach is used to compute the safety score for each instruction. We have evaluated the approach through ABB's RobotStudio®simulator. We simulate two scenarios: First, where two robots share a workspace and second, where a robot moves along a path with an obstacle. The simulations show that using our code analyzer and safety score formalism; it is possible to evaluate the application code for safety and enable avoidance of potentially unsafe behavior.
Avijit Mandal, Divyasheel Sharma, Mohak Sukhwani, Raoul Praful Jetley, Santonu Sarkar
INDIN5
2018 Modeling Operational Fairness of Hybrid Cloud Brokerage
abstract
Cloud service brokerage is an emerging technology that attempts to simplify the consumption and operation of hybrid clouds. Today's cloud brokers attempt to insulate consumers from the vagaries of multiple clouds. To achieve the insulation, the modern cloud broker needs to disguise itself as the end-provider to consumers by creating and operating a virtual data center construct that we call a "meta-cloud", which is assembled on top of a set of participating supplier clouds. It is crucial for such a cloud broker to be considered a trusted partner both by cloud consumers and by the underpinning cloud suppliers. A fundamental tenet of brokerage trust is vendor neutrality. On the one hand, cloud consumers will be comfortable if a cloud broker guarantees that they will not be led through a preferred path. And on the other hand, cloud suppliers would be more interested in partnering with a cloud broker who promises a fair apportioning of client provisioning requests. Because consumer and supplier trust on a meta-cloud broker stems from the assumption of being agnostic to supplier clouds, there is a need for a test strategy that verifies the fairness of cloud brokerage. In this paper, we propose a calculus of fairness that defines the rules to determine the operational behavior of a cloud broker. The calculus uses temporal logic to model the fact that fairness is a trait that has to be ascertained over time; it is not a characteristic that can be judged at a per-request fulfillment level. Using our temporal calculus of fairness as the basis, we propose an algorithm to determine the fairness of a broker probabilistically, based on its observed request apportioning policies. Our model for the fairness of cloud broker behavior also factors in inter-provider variables such as cost divergence and capacity variance. We empirically validate our approach by constructing a meta-cloud from AWS, Azure and IBM, in addition to leveraging a cloud simulator. Our industrial engagements with large enterprises also validate the need for such cloud brokerage with verifiable fairness.
Sreekrishnan Venkateswaran, Santonu Sarkar
CCGrid2
2018 Towards Transforming an Industrial Automation System from Monolithic to Microservices
abstract
Container technology enables designers to build (micro)service-oriented systems with on-demand scalability and availability easily, provided the original system has been well-modularized to begin with. Industry automation applications, built a long time ago, aim to adopt this technology to become more flexible and ready to be a part of the internet of thing based next-generation industrial system. In this paper, we share our work-in-progress experience of transforming a complex, distributed industrial automation system to a microservice based containerized architecture. We propose a containerized architecture of the “to-be” system and observe that despite being distributed, the “as-is” system tend to follow a monolithic architecture with strong coupling among the participating components. Consequently it becomes difficult to achieve the proposed microservice based architecture without a significant change. We also discuss the workload handling, resource utilization and reliability aspects of the “to-be” architecture using a prototype implementation.
Santonu Sarkar, Gloria Vashi, P. P. Abdulla
ETFA1
2018 Analysis of GPGPU Programs for Data-race and Barrier Divergence
Santonu Sarkar, Prateek Kandelwal, Soumyadip Bandyopadhyay, Holger Giese
ICSOFT1
2018 Thrust2D: A new design abstraction framework for structured grid class of algorithms
abstract
Summary An important goal of structured parallel programming has been to provide a design framework that balances between the extent of abstraction built over the hardware and the amount of control given to the programmer to leverage the hardware resource features. Towards this goal, NVIDIA has released an open‐source design framework called Thrust based on C++ STL, where the developers can express the functionality in STL style, without having to know the architectural details of the underlying parallel infrastructure. While the framework is generic and portable, it does not support the right abstraction for two‐dimensional data, which is heavily used in most of the popular parallel algorithms. In this paper, we proposed Thrust2D, an extension of Thrust to support the abstraction for two‐dimensional data, targeted towards structured grid class of applications. We took several structured grid examples from Rodinia benchmark, OpenCV framework, and NVIDIA samples and rewrote them using Thrust2D. We demonstrated that, in some cases, we get nearly 80% reduction in code complexity, and for 12 out of 17 applications we have tested, the kernel performance of Thrust2D versions are well within 85% of the native CUDA versions. When we consider the total execution time, 14 out of 17 Thrust2D versions performance are within 85% of the native CUDA versions. In some cases, the performance of the Thrust2D versions has outperformed the native versions.
Santonu Sarkar, Ajai V. George, Sankar Manoj
Concurr. Comput. Pract. Exp.1
2018 Optimizing MapReduce for energy efficiency
abstract
Summary The efficient use of energy is essential to address concerns of cost and sustainability. Many data centers contain MapReduce clusters to process Big Data applications. A large number of machines and fault tolerance capabilities make MapReduce clusters energy inefficient. In this paper, we present a Configurator based on performance and energy models to improve the energy efficiency of MapReduce systems. Our solution is novel as it takes into account the dependence of the performance and energy consumption of a cluster on MapReduce parameters. While this dependence is known, we are the first to model it and design a Configurator to optimize these parameter settings for maximizing the energy efficiency of MapReduce systems. Our empirical evaluations show that the Configurator can result in up to 50% improvement in the energy efficiency of typical MapReduce applications in two architecturally different clusters.
Nidhi Tiwari, Umesh Bellur, Santonu Sarkar, Maria Indrawan
Softw. Pract. Exp.3
2018 Architectural partitioning and deployment modeling on hybrid clouds
abstract
Summary The hybrid cloud idea is increasingly gaining momentum because it brings distinct advantages as a hosting platform for complex software systems. However, there are several challenges that need to be surmounted before hybrid hosting can become pervasive and penetrative. One main problem is to architecturally partition workloads across permutations of feasible cloud and non‐cloud deployment choices to yield the best‐fit hosting combination. Another is to predict the effort estimate to deliver such an advantageous hybrid deployment. In this paper, we describe a heuristic solution to address the said obstacles and converge on the ideal hybrid cloud deployment architecture, based on properties and characteristics of workloads that are sought to be hosted. We next propose a model to represent such a hybrid cloud deployment and demonstrate a method to estimate the effort required to implement and sustain that deployment. We also validate our model through dozens of case studies spanning several industry verticals and record results pertaining to how the industrial grouping of a software system can impact the aforementioned hybrid deployment model. Copyright © 2017 John Wiley & Sons, Ltd.
Sreekrishnan Venkateswaran, Santonu Sarkar
Softw. Pract. Exp.2
2017 SamaTulyata: An Efficient Path Based Equivalence Checking Tool
Soumyadip Bandyopadhyay, Santonu Sarkar, Dipankar Sarkar 0001, Chittaranjan A. Mandal
ATVA2
2017 Thrust++: Extending Thrust Framework for Better Abstraction and Performance
abstract
A good design abstraction framework for high performance computing should provide a higher level programming abstraction that strikes a balance between the abstraction and visibility over the hardware so that the software developer can write a portable software without having to understand the hardware nuances, yet exploit the compute power optimally. In this paper we have analyzed a popular design abstraction framework called "Thrust" from NVIDIA, and proposed an extension called Thrust++ that provides abstraction over the memory hierarchy of an NVIDIA GPU. Thrust++ allows developers to make efficient use of shared memory and overall, provides better control over the GPU memory hierarchy while writing applications in Thrust style for the CUDA backend. We have shown that when applications are written for the CUDA backend using Thrust++, they have minimal performance degradation when compared to their equivalent CUDA versions. Further, Thrust++ provides almost 4x speedup when compared to Thrust, for certain compute intensive kernels that repeatedly use the reduce operation.
Ajai V. George, Sankar Manoj, Sanket Rajan Gupte, Sayantan Mitra, Santonu Sarkar
HiPC5
2017 An End-to-end Formal Verifier for Parallel Programs
Soumyadip Bandyopadhyay, Santonu Sarkar, Kunal Banerjee 0001
ICSOFT2
2017 Analysis and Diagnosis of SLA Violations in a Production SaaS Cloud
abstract
A software-as-a-service (SaaS) needs to provide its intended service as per its stated service-level agreements (SLAs). While SLA violations in a SaaS platform have been reported, not much work has been done to empirically characterize failures of SaaS. In this paper, we study SLA violations of a production SaaS platform, diagnose the causes, unearth several critical failure modes, and then, suggest various solution approaches to increase the availability of the platform as perceived by the end user. Our approach combines field failure data analysis (FFDA) and fault injection. Our study is based on 283 days of operational logs of the platform. During this time, the platform received business workload from 42 customers spread over 22 countries. We have first developed a set of home-grown FFDA tools to analyze the log, and second implemented a fault injector to automatically inject several runtime errors in the application code written in .NET/C#, and then, collate the injection results. We summarize our finding as: first, system failures have caused 93% of all SLA violations; second, our fault injector has been able to recreate a few cases of bursts of SLA violations that could not be diagnosed from the logs; and third, the fault injection mechanism could recreate several error propagation paths leading to data corruptions that the failure data analysis could not reveal. Finally, the paper presents some system-level implication of this study and how the joint use of fault injection and log analysis may help in improving the reliability of the measured platform.
Catello Di Martino, Santonu Sarkar, Rajeshwari Ganesan, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer
IEEE Trans. Reliab.2
2016 CPU Frequency Tuning to Improve Energy Efficiency of MapReduce Systems
abstract
Energy efficiency is a major concern in today's data centers that house large scale distributed processing systems such as data parallel MapReduce clusters. Modern power aware systems utilize the dynamic voltage and frequency scaling mechanism available in processors to manage the energy consumption. In this paper, we initially characterize the energy efficiency of MapReduce jobs with respect to built-in power governors. Our analysis indicates that while a built-in power governor provides the best energy efficiency for a job that is CPU as well as IO intensive, a common CPU-frequency across the cluster provides best the energy efficiency for other types of jobs. In order to identify this optimal frequency setting, we derive energy and performance models for MapReduce jobs on a HPC cluster and validate these models experimentally on different platforms. We demonstrate how these models can be used to improve energy efficiency of the machine learning MapReduce applications running on the Yarn platform. The execution of jobs at their optimal frequencies improves the energy efficiency by average 25% over the default governor setting. In case of mixed workloads, the energy efficiency improves by up to 10% when we use an optimal CPU-frequency across the cluster.
Nidhi Tiwari, Umesh Bellur, Santonu Sarkar, Maria Indrawan
ICPADS3
2016 Unified power and energy measurement API for HPC co-processors
abstract
Power and energy optimization of applications running on a High Performance Computing infrastructure is an important research area in today's HPC driven world. Considerable amount of investments are being made for the same, emphasizing on the importance of such a research. The above mentioned optimization is a two step process. First, one must measure and analyze the power consumption accurately and then optimize the application by modifying the relevant sections of the application software responsible for it. In this paper we propose an API, called the Unified Power Profiling API (UPPAPI), which allows a parallel application developer to measure the power and energy consumed by the heterogeneous system during the application execution. The first important feature of this API is the abstraction created over different co-processors. Second, we have incorporated a corrective power model to remove the discrepancy arising out of the on-board GPU sensors. We illustrate the working of the above mentioned tool using Rodinia and SHOC Benchmarks. We further verify the correctness of the API generated measurement by comparing the results returned by the API with the actual power measured with a Krykard ALM10 power analyzer. This paper is an early attempt to develop a well rounded software tool, which enables the developer to optimize the resource usage and create efficient applications.
Mohak Chadha, Abhishek Srivastava 0001, Santonu Sarkar
IPCCC3
2016 How Long Will This Live? Discovering the Lifespans of Software Engineering Ideas
abstract
We all want to be associated with long lasting ideas; as originators, or at least, expositors. For a tyro researcher or a seasoned veteran, knowing how long an idea will remain interesting in the community is critical in choosing and pursuing research threads. In the physical sciences, the notion of half-life is often evoked to quantify decaying intensity. In this paper, we study a corpus of 19,000+ papers written by 21,000+ authors across 16 software engineering publication venues from 1975 to 2010, to empirically determine the half-life of software engineering research topics. In the absence of any consistent and well-accepted methodology for associating research topics to a publication, we have used natural language processing techniques to semi-automatically identify and associate a set of topics with a paper. We adapted measures of half-life already existing in the bibliometric context for our study, and also defined a new measure based on publication and citation counts. We find evidence that some of the identified research topics show a mean half-life of close to 15 years, and there are topics with sustaining interest in the community. We report the methodology of our study in this paper, as well as the implications and utility of our results.
Subhajit Datta, Santonu Sarkar, A. S. M. Sajeev
IEEE Trans. Big Data2
2015 The Importance of Being Isolated: An Empirical Study on Chromium Reviews
abstract
As large scale software development has become more collaborative, and software teams more globally distributed, several studies have explored how developer interaction influences software development outcomes. The emphasis so far has been largely on outcomes like defect count, the time to close modification requests etc. In the paper, we examine data from the Chromium project to understand how different aspects of developer discussion relate to the closure time of reviews. On the basis of analyzing reviews discussed by 2000+ developers, our results indicate that quicker closure of reviews owned by a developer relates to higher reception of information and insights from peers. However, we also find evidence that higher engagement in collaboration by a developer is associated with slower closure of the reviews she owns. Within the scope of our study, these results lead us to conclude that peer review of code may have a distinct dynamic that is facilitated by developers working in relative isolation.
Subhajit Datta, Devarshi Bhatt, Proshanta Sarkar, Santonu Sarkar
ESEM5
2013 Factors Influencing Research Contributions and Researcher Interactions in Software Engineering: An Empirical Study
abstract
Research into software engineering (SE) education is largely concentrated on teaching and learning issues in coursework programs. This paper, in contrast, provides a meta analysis of research publications in software engineering to help with research education in SE. Studying publication patterns in a discipline will assist research students and supervisors gain a deeper understanding of how successful research has occurred in the discipline. We present results from a large scale empirical study covering over three and a half decades of software engineering research publications. We identify how different factors of publishing relate to the number of papers published as well as citations received for a researcher, and how the most successful researchers collaborate and co-cite one another. Our results show that authors with high publication rates do not concentrate on a few selected venues to publish, researchers with high publication rates behave differently from researchers of high citation rates (with the latter group co-authoring and citing their peers to a much lesser extent than the former), and collaborators citing each other's works is not a significant phenomenon in SE research.
Subhajit Datta, A. S. M. Sajeev, Santonu Sarkar, Nishant Kumar 0002
APSEC (1)3
2013 Keep it moving: Proactive workload management for reducing SLA violations in large scale SaaS clouds
abstract
Software failures, workload-related failures and job overload conditions bring about SLA violations in software-as-a-service (SaaS) systems. Existing work does not address mitigation of SLA violations completely as (i) none of them address mitigation of SLA violations in business specific scenarios (SaaS, in our case), (ii) while some do not address software and workload-related failures, other approaches do not address the problem of target PM selection for workload migration comprehensively (leaving out vital considerations like workload compatibility checks between migrating VM and VMs at the target PM) and (iii) a clear mathematical mapping between workload, resource demand and SLA is lacking. In this paper, we present the Keep It Moving (KIM) software framework for the cloud controller that helps minimize service failures due to SLA violation of availability, utilization and response time in SaaS cloud data centers. Though we consider migration to be the primary mitigation technique, we also try to mitigate SLA violations without migration. We achieve this by performing a capacity check on the host physical machine (PM) before the migration to identify if enough capacity is available on the current PM to address the upcoming SLA violations by restart/reboot or VM resizing. In certain cases such as workload-related failures due to corrupt files, we prefer workload rerouting to a replica VM over migration. We formulate the selection of a target PM as a multi-objective optimization problem. We validate our proposed approach by using a trace-based discrete event simulation of a virtualized data center where failure and workload characteristics are simulated from data extracted from a real SaaS business server logs. We found that a 60% reduction in SLA violation is possible using our approach as well as reducing VM downtime by approximately 10%.
Arpan Roy, Rajeshwari Ganesan, Santonu Sarkar
ISSRE3
2012 Analysis of SaaS Business Platform Workloads for Sizing and Collocation
abstract
Sharing of physical infrastructure using virtualization presents an opportunity to improve the overall resource utilization. It is extremely important for a Software as a Service (SaaS) provider to understand the characteristics of the business application workload in order to size and place the virtual machine (VM) containing the application. A typical business application has a multi-tier architecture and the application workload is often predictable. Using the knowledge of the application architecture and statistical analysis of the workload, one can obtain an appropriate capacity and a good placement strategy for the corresponding VM. In this paper we propose a tool iCirrus-WoP that determines VM capacity and VM collocation possibilities for a given set of application workloads. We perform an empirical analysis of the approach on a set of business application workloads obtained from geographically distributed data centers. The iCirrus-WoP tool determines the fixed reserved capacity and a shared capacity of a VM which it can share with another collocated VM. Based on the workload variation, the tool determines if the VM should be statically allocated or needs a dynamic placement. To determine the collocation possibility, iCirrus-WoP performs a peak utilization analysis of the workloads. The empirical analysis reveals the possibility of collocating applications running in different time-zones. The VM capacity that the tool recommends, show a possibility of improving the overall utilization of the infrastructure by more than 70% if they are appropriately collocated.
Rajeshwari Ganesan, Santonu Sarkar, Akshay Narayan 0002
IEEE CLOUD2
2012 iCirrus Wop: Workload Analysis for Virtual Machine Placements
abstract
True essence of the technology of virtualization is the ability to allow one or more workloads to share the underlying physical resources, thereby bringing about significant cost saving. However, in order to maximize the cost savings from this disruptive technology, it is essential to adopt optimal resource management techniques. These techniques broadly encompass approaches to virtual machine (VM) sizing and placement in a manner that maximizes the physical infrastructure utilization, alongside ensuring that the desired service-level objectives of the candidate workloads are met. In this paper, we propose a novel workload analysis approach for VM placement, which relies on examining the time varying processing demands and variability of the workloads to determine the most optimal placement. Such a solution will result in maximizing infrastructure utilization and ensure that the SLAs of the candidate workloads are met after placement. The technique has been effectively applied to real-life workloads that pertain to SaaS based business platforms offered to clients spread across different geographical locations. A paper based assessment reported over 25% improvement in the overall infrastructure utilization by using the proposed algorithm as compared to other well-known approaches.
Geetika Goel, Rajeshwari Ganesan, Santonu Sarkar, Kavish Kaup
ICPADS3
2012 Reuse and Refactoring of GPU Kernels to Design Complex Applications
abstract
Developers of GPU kernels, such as FFT, linear solvers, etc, tune their code extensively in order to obtain optimal performance, making efficient use of different resources available on the GPU. Complex applications are composed of several such kernel components. The software engineering community has performed extensive research on component based design to build generic and flexible components, such that a component can be reused across diverse applications, rather than optimizing its performance. Since a GPU is used primarily to improve performance, application performance becomes a key design issue. The contribution of our work lies in extending component based design research in a new direction, dealing with the performance impact of refactoring an application consisting of the composition of highly tuned kernels. Such refactoring can make the composition more effective with respect to GPU resource usage especially when combined with suitable scheduling. Here we propose a methodology where developers of highly tuned kernels can enable application designers to optimize performance of the composition. Kernel developers characterize the performance of a kernel through its "performance signature". The application designer combines these kernels such that the performance of the refactored kernel is better than the sum of the performances of the individual kernels. This is partly based on the observation that different kernels may make unbalanced use of different GPU resources like different types of memory. Kernels may also have the potential to share data. Refactoring the kernels, combining them, and scheduling them suitably can improve performance. We study different types of potential design optimizations and evaluate their effectiveness on different types of kernels. This may even involve choosing non-optimal parameters for an individual kernel. We analyze how the performance signature of the composition changes from that of the individual kernels through our techniques. We demonstrate that our techniques lead to over 50% improvement with some kernels. Furthermore, the performance of a basic molecular dynamics application can be improved by around 25.7%, on a Fermi GPU, compared with an un-refactored implementation.
Santonu Sarkar, Sayantan Mitra, Ashok Srinivasan
ISPA1
2011 Implementation of a Scalable Next Generation Sequencing Business Cloud Platform-An Experience Report
abstract
Life science industry is looking towards new and cost-effective ways to manage and analyze huge amount of genomic data for faster innovation in drug or biologics discovery. To that effect, various alliances among competitive organizations are getting formed, such as the Pistoia Alliance, to collaborate and share a pool of genomic data and build useful search and analysis techniques for the alliance partners. In order to make the development, and management of data and applications cost-effective, a secure cloud computing based platforms are being considered. In this paper we describe an experience report of building such a collaborative platform on Amazon cloud platform. In order to build a scalable genome sequence alignment solution, we have adopted the well-known BLAST framework on Hadoop platform. A major challenge here is that the BLAST executable requires to be ported as it is, and yet the execution needs to scale, as the number of jobs increases, by elastically growing the Hadoop infrastructure. In this paper we proposed a BLAST database partitioning solution to achieve optimal scalability. Our controlled experiment is encouraging, the empirical result shows that the job execution scales with the number of jobs, if the partition sizes are chosen appropriately.
Shyam Kumar Doddavula, Madhavi Rani, Santonu Sarkar, Harsh Rajesh Vachhani, Akansha Jain, Mudit Kaushik
IEEE CLOUD3
2009 Extracting High-Level Functional Design from Software Requirements
abstract
Practitioners spend significant amounts of time creating high-level design from requirements. Though there exist methodologies to describe and manage requirements and design artifacts, there is not yet an automated way to faithfully translate a requirement into a high-level design. While it is extremely difficult to generate design elements from free-form natural language due to its inherent ambiguity, it is possible to significantly improve the accuracy of the design from relatively structured and constrained natural language. In this paper we propose a technique to generate high-level class diagrams from a set of requirements, using a set of requirement-specific heuristics. In this approach, we leverage work we had previously done to first process a requirement statement to classify it into a requirement type, and then break it into various constituents. Depending on the requirement type and its constituents, our heuristics then discover a functional design comprising of coarse-grained modules, their relationships and responsibilities. We express the design as a UML class diagram in IBM rational software architect (RSA) format. Our preliminary investigation shows that the resulting class diagram is rich, and can be used by practitioners as a basis for further design.
Vibhu Saujanya Sharma, Santonu Sarkar, Kunal Verma, Arun Panayappan, Alex Kass
APSEC2
2009 Discovery of architectural layers and measurement of layering violations in source code
Santonu Sarkar, Girish Maskeri Rama, Shubha Ramachandran
J. Syst. Softw.1
2008 Metrics for Measuring the Quality of Modularization of Large-Scale Object-Oriented Software
abstract
The metrics formulated to date for characterizing the modularization quality of object-oriented software have considered module and class to be synonymous concepts. But a typical class in object oriented programming exists at too low a level of granularity in large object-oriented software consisting of millions of lines of code. A typical module (sometimes referred to as a superpackage) in a large object-oriented software system will typically consist of a large number of classes. Even when the access discipline encoded in each class makes for "clean" class-level partitioning of the code, the intermodule dependencies created by associational, inheritance-based, and method invocations may still make it difficult to maintain and extend the software. The goal of this paper is to provide a set of metrics that characterize large object-oriented software systems with regard to such dependencies. Our metrics characterize the quality of modularization with respect to the APIs of the modules, on the one hand, and, on the other, with respect to such object-oriented inter-module dependencies as caused by inheritance, associational relationships, state access violations, fragile base-class design, etc. Using a two-pronged approach, we validate the metrics by applying them to popular open-source software systems.
Santonu Sarkar, Avinash C. Kak, Girish Maskeri Rama
IEEE Trans. Software Eng.1
2007 API-Based and Information-Theoretic Metrics for Measuring the Quality of Software Modularization
abstract
We present in this paper a new set of metrics that measure the quality of modularization of a non-object-oriented software system. We have proposed a set of design principles to capture the notion of modularity and defined metrics centered around these principles. These metrics characterize the software from a variety of perspectives: structural, architectural, and notions such as the similarity of purpose and commonality of goals. (By structural, we are referring to intermodule coupling-based notions, and by architectural, we mean the horizontal layering of modules in large software systems.) We employ the notion of API (application programming interface) as the basis for our structural metrics. The rest of the metrics we present are in support of those that are based on API. Some of the important support metrics include those that characterize each module on the basis of the similarity of purpose of the services offered by the module. These metrics are based on information-theoretic principles. We tested our metrics on some popular open-source systems and some large legacy-code business applications. To validate the metrics, we compared the results obtained on human-modularized versions of the software (as created by the developers of the software) with those obtained on randomized versions of the code. For randomized versions, the assignment of the individual functions to modules was randomized
Santonu Sarkar, Girish Maskeri Rama, Avinash C. Kak
IEEE Trans. Software Eng.1
2006 A Method for Detecting and Measuring Architectural Layering Violations in Source Code
abstract
The layered architecture pattern has been widely adopted by the developer community in order to build large software systems. The layered organization of software modules offers a number of benefits such as reusability, changeability and portability to those who are involved in the development and maintenance of such software systems. But in reality as the system evolves over time, rarely does the actual source code of the system conform to the conceptual horizontal layering of modules. This in turn results in a significant degradation of system maintainability. In order to re-factor such a system to improve its maintainability, it is very important to discover, analyze and measure violations of layered architecture pattern. In this paper we propose a technique to discover such violations in the source code and quantitatively measure the amount of non-conformance to the conceptual layering. The proposed approach evaluates the extent to which the module dependencies across layers violate the layered architecture pattern. In order to evaluate the accuracy of our approach, we have applied this technique to discover and analyze such violations to a set of open source applications and a proprietary business application by taking the help of domain experts wherever possible.
Santonu Sarkar, Girish Maskeri Rama, Shubha Ramachandran
APSEC1
2005 Metrics for Analyzing Module Interactions in Large Software Systems
abstract
We present a new set of metrics for analyzing the interaction between the modules of a large software system. We believe that these metrics would be important to any automatic or semi-automatic code modularization algorithm. The metrics are based on the rationale that code partitioning should be based on the principle of similarity of service provided by the different functions encapsulated in a module. Although module interaction metrics are necessary for code modularization, in practice they must be accompanied by metrics that measure other important attributes of how the code is partitioned into modules. These other metrics, dealing with code properties such as the approximate uniformity of module sizes, conformance to any size constraints on the modules, etc., are also included in the work presented here. To give the reader some insight into the workings of our metrics, this paper also includes some results obtained by applying the metrics to the body of code that constitutes the open-source Apache HTTP server. We apply our metrics to this code as packaged by the developers of the software and to the other partially and fully randomized versions of the code.
Santonu Sarkar, Avinash C. Kak, N. S. Nagaraja
APSEC1
1994 Interface design and controller synthesis of digital systems in an object oriented environment
Santonu Sarkar, Arun K. Majumdar, Anupam Basu
Microprocess. Microprogramming1
1991 VLODS: a VLSI object oriented database system
Tapas K. Nayak, Arun K. Majumdar, Anupam Basu, Santonu Sarkar
Inf. Syst.4