K. Gopinath

dblp:48/2735 · also Kanchi Gopinath · DBLP profile ↗
← Back
45ranked-venue papers
5as first author
3since 2021 · last 2025
0000-0003-4478-0064ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 23 · 4 first-author · 1 since 2021Software engineering, systems software and programming languages · 16 · 1 first-author · 2 since 2021Theory of computation · 4Applied, interdisciplinary, general and emerging computing · 4 · 1 since 2021Artificial intelligence and machine learning · 3 · 1 since 2021Security and privacy · 1Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2025 EMD: Fair and Efficient Dynamic Memory De-bloating of Transparent Huge Pages
abstract
Recent processors rely on huge pages to reduce the cost of virtual-to-physical address translation. However, huge pages are notorious for creating memory bloat – a phenomenon wherein the OS ends up allocating more physical memory to an application than its actual requirement. This extra memory can be reclaimed by the OS via de-bloating at runtime. However, we find that current OS-level solutions either lack support for dynamic memory de-bloating, or suffer from performance and fairness pathologies while de-bloating. We address these issues with EMD (Efficient Memory De-bloating). The key insight in EMD is that different regions in an application’s address space exhibit different amounts of memory bloat. Consequently, the tradeoff between memory efficiency and performance varies significantly within a given application e.g., we find that memory bloat is typically concentrated in specific regions, and de-bloating them leads to minimal performance impact. Hinged on this insight, EMD employs a prioritization scheme for fine-grained, efficient, and fair reclamation of memory bloat. EMD improves performance by up to 69% compared to HawkEye — a state-of-the-art OS-based huge page management system. EMD also eliminates fairness concerns associated with dynamic memory de-bloating.
Parth Gangar, Ashish Panwar, K. Gopinath
ISMM3
2022 Enhancing the Credit Card Fraud Detection Using Decision Tree and Adaptive Boosting Techniques
K. R. Prasanna Kumar, S. Aravind, K. Gopinath, P. Navienkumar, K. Logeswaran 0001, M. Gunasekar
ISDA (3)3
2021 Fast local page-tables for virtualized NUMA servers with vMitosis
abstract
Increasing heterogeneity in the memory system mandates careful data placement to hide the non-uniform memory access (NUMA) effects on applications. However, NUMA optimizations have predominantly focused on application data in the past decades, largely ignoring the placement of kernel data structures due to their small memory footprint; this is evident in typical OS designs that pin kernel objects in memory. In this paper, we show that careful placement of kernel data structures is gaining importance in the context of page-tables: sub-optimal placement of page-tables causes severe slowdown (up to 3.1x) on virtualized NUMA servers.
Ashish Panwar, Reto Achermann, Arkaprava Basu, Abhishek Bhattacharjee, K. Gopinath, Jayneel Gandhi
ASPLOS5
2020 Need for a Deeper Cross-Layer Optimization for Dense NAND SSD to Improve Read Performance of Big Data Applications: A Case for Melded Pages
Arpith K, K. Gopinath
HotStorage2
2019 HawkEye: Efficient Fine-grained OS Support for Huge Pages
abstract
Effective huge page management in operating systems is necessary for mitigation of address translation overheads. However, this continues to remain a difficult area in OS design. Recent work on Ingens uncovered some interesting pitfalls in current huge page management strategies. Using both page access patterns discovered by the OS kernel and fine-grained data from hardware performance counters, we expose problematic aspects of current huge page management strategies. In our system, called HawkEye/Linux, we demonstrate alternate ways to address issues related to performance, page fault latency and memory bloat; the primary ideas behind HawkEye management algorithms are async page pre-zeroing, de-duplication of zero-filled pages, fine-grained page access tracking and measurement of address translation overheads through hardware performance counters. Our evaluation shows that HawkEye is more performant, robust and better-suited to handle diverse workloads when compared with current state-of-the-art systems.
Ashish Panwar, Sorav Bansal, K. Gopinath
ASPLOS3
2018 Making Huge Pages Actually Useful
abstract
The virtual-to-physical address translation overhead, a major performance bottleneck for modern workloads, can be effectively alleviated with huge pages. However, since huge pages must be mapped contiguously, OSs have not been able to use them well because of the memory fragmentation problem despite hardware support for huge pages being available for nearly two decades. This paper presents a comprehensive study of the interaction of fragmentation with huge pages in the Linux kernel. We observe that when huge pages are used, problems such as high CPU utilization and latency spikes occur because of unnecessary work (e.g., useless page migration) performed by memory management related subsystems due to the poor handling of unmovable (i.e., kernel) pages. This behavior is even more harmful in virtualized systems where unnecessary work may be performed in both guest and host OSs. We present Illuminator, an efficient memory manager that provides various subsystems, such as the page allocator, the ability to track all unmovable pages. It allows subsystems to make informed decisions and eliminate unnecessary work which in turn leads to cost-effective huge page allocations. Illuminator reduces the cost of compaction (up to 99%), improves application performance (up to 2.3x) and reduces the maximum latency of MySQL database server (by 30x). Importantly, this work shows the effectiveness of a simple solution for long-standing huge page related problems.
Ashish Panwar, Aravinda Prasad, K. Gopinath
ASPLOS3
2018 A frugal approach to reduce RCU grace period overhead
abstract
Grace period computation is a core part of the Read-Copy-Update (RCU) synchronization technique that determines the safe time to reclaim the deferred objects' memory. We first show that the eager grace period computation employed in the Linux kernel is appropriate only for enterprise workloads such as web and database servers where a large amount of reclaimable memory awaits the completion of a grace period. However, such memory is negligible in High-Performance Computing (HPC) and mostly idling environments due to limited OS kernel activity. Hence an eager approach is not only futile but also detrimental as the CPU cycles consumed to compute a grace period leads to jitter in HPC and frequent CPU wake-ups in idle environments.
Aravinda Prasad, K. Gopinath
EuroSys2
2018 Probabilistic Sequential Consistency in Social Networks
abstract
Researchers have proposed numerous consistency models in distributed systems that offer higher performance than classical sequential consistency (SC). Even though these models do not guarantee sequential consistency; they either behave like an SC model under certain restrictive scenarios, or ensure SC behavior for a part of the system. We propose a different line of thinking where we try to accurately estimate the number of SC violations, and then try to adapt our system to optimally tradeoff performance, resource usage, and the number of SC violations. In this paper, we propose a generic theoretical model that can be used to analyze systems that are comprised of multiple sub-domains - each sequentially consistent. It is validated with real world measurements. Next, we use this model to propose a new form of consistency called social consistency, where socially connected users perceive an SC execution, whereas the rest of the users need not. We create a prototype social network application and implement it on the Cassandra key-value store. We show that our system has 2.4× more throughput than Cassandra and provides 37% better quality-of-experience.
Priyanka Singla 0001, Shubhankar Suman Singh, K. Gopinath, Smruti R. Sarangi
HiPC3
2017 Scalable Performance Tuning of Hadoop MapReduce: A Noisy Gradient Approach
abstract
Hadoop MapReduce is a popular framework for distributed storage and processing of large datasets and is used for big data analytics. It has various configuration parameters which play an important role in deciding the performance i.e., the execution time of a given big data processing job. Default values of these parameters do not result in good performance and therefore it is important to tune them. However, there is inherent difficulty in tuning the parameters due to two important reasons - first, the parameter search space is large and second, there are cross-parameter interactions. Hence, there is a need for a dimensionality-free method which can automatically tune the configuration parameters by taking into account the cross-parameter dependencies. In this paper, we propose a novel Hadoop parameter tuning methodology, based on a noisy gradient algorithm known as the simultaneous perturbation stochastic approximation (SPSA). The SPSA algorithm tunes the selected parameters by directly observing the performance of the Hadoop MapReduce system. The approach followed is independent of parameter dimensions and requires only 2 observations per iteration while tuning. We demonstrate the effectiveness of our methodology in achieving good performance on popular Hadoop benchmarks namely Grep, Bigram, Inverted Index, Word Co-occurrence and Terasort. Our method, when tested on a 25 node Hadoop cluster shows 45-66% decrease in execution time of Hadoop jobs on an average, when compared to prior methods. Further, our experiments also indicate that the parameters tuned by our method are resilient to changes in number of cluster nodes, which makes our method suitable to optimize Hadoop when it is provided as a service on the cloud.
Sindhu Padakandla, Chandrashekar Lakshminarayanan, Priyank Parihar, K. Gopinath, Shalabh Bhatnagar
CLOUD5
2017 The RCU-Reader Preemption Problem in VMs
Aravinda Prasad, K. Gopinath, Paul E. McKenney
USENIX ATC2
2016 Prudent Memory Reclamation in Procrastination-Based Synchronization
abstract
Procrastination is the fundamental technique used in synchronization mechanisms such as Read-Copy-Update (RCU) where writers, in order to synchronize with readers, defer the freeing of an object until there are no readers referring to the object. The synchronization mechanism determines when the deferred object is safe to reclaim and when it is actually reclaimed. Hence, such memory reclamations are completely oblivious of the memory allocator state. This induces poor memory allocator performance, for instance, when the reclamations are ill-timed. Furthermore, deferred objects provide hints about the future that inform memory regions that are about to be freed. Although useful, hints are not exploited as deferred objects are not visible to memory allocators. We introduce Prudence, a dynamic memory allocator, that is tightly integrated with the synchronization mechanism to ensure visibility of deferred objects to the memory allocator. Such an integration enables Prudence to (i) identify the safe time to reclaim deferred objects' memory, (ii) have an inclusive view of the allocated, free and about-to-be-freed objects, and (iii) exploit optimizations based on the hints about the future during important state transitions. Our evaluation in the Linux kernel shows that Prudence integrated with RCU performs 3.9X to 28X better in micro-benchmarks compared to SLUB, a recent memory allocator in the Linux kernel. It also improves the overall performance perceptibly (4%-18%) for a mix of widely used synthetic and application benchmarks. Further, it performs better (up to 98%) in terms of object hits in caches, object cache churns, slab churns, peak memory usage and total fragmentation, when compared with the SLUB allocator.
Aravinda Prasad, K. Gopinath
ASPLOS2
2016 A Secure Role-Based Cloud Storage System For Encrypted Patient-Centric Health Records
abstract
With the rapid developments occurring in cloud services, there has been a growing trend to use cloud for large-scale data storage. Due to the increasing popularity of cloud storage, many healthcare organizations have started moving electronic health records (EHRs) to cloud-based storage systems. However, this has raised the important security issue of how to protect and prevent unauthorized access to EHR data stored in a public cloud. Several cryptographic access control schemes have been proposed to protect the security of data stored in the cloud by integrating cryptographic techniques with access control models. In this paper, we consider a novel role-based encryption technique to build a secure and flexible large-scale EHR system where role-based access control policies are enforced in a cloud environment. Then we discuss a practical EHR system called the personally controlled electronic health record (PCEHR) system recently developed by the Australian Government, and show how the security weaknesses in the PCEHR system can be addressed by our proposed scheme. The proposed system has the potential to be useful in commercial healthcare systems as it captures practical access policies based on roles in a flexible manner and provides secure data storage in the cloud enforcing these access policies.
Vijay Varadharajan, K. Gopinath
Comput. J.3
2015 Towards Practical Page Placement for a Green Memory Manager
abstract
Increased performance demand of modern applications has resulted in large memory modules and higher performance processors in computing systems. Power consumption becomes an important aspect when these resources go underutilized in a running system, e.g. during idle periods or lighter workloads. CPUs have come a long way in optimizing away the unnecessary power consumption in both hardware and software for such scenarios through solutions like Dynamic Voltage/Frequency Scaling. However, support for memory power optimization is still missing in modern operating systems despite hardware support being available for many years in the form of multiple power states and techniques like Partial Array Self-Refresh. In this work, we explore the behavior of Linux memory manager and report that even at 10% of memory utilization, there are references to all physical memory banks in a long running system due to random page allocation and ignorance of memory bank boundaries. These references can be consolidated to a subset of memory banks by using page migration techniques. Unfortunately, migration of large contiguous blocks is often restricted due to the presence of unmovable pages primarily owned by kernel. We provide some techniques for utilizing the hardware facilitated Partial Array Self-Refresh by introducing bank awareness in the existing buddy allocation framework of Linux memory manager as well as for improving the page migration support of large contiguous blocks. Through a set of simple changes in Linux VM, we have been able to reduce the number of referenced memory banks significantly. Memory-hotplug framework, which relies on page migration of large contiguous blocks, also shows significant improvement in terms of number of removable memory sections. Benchmark results show no performance degradation in the modified kernel which makes the proposed solution desirable.
Ashish Panwar, K. Gopinath
HiPC2
2014 Human Activity Recognition by Matching Curve Shapes
Poorna Talkad Sukumar, K. Gopinath
ICONIP (2)2
2013 Elastic Resources Framework in IaaS, Preserving Performance SLAs
abstract
Elasticity in cloud systems provides the flexibility to acquire and relinquish computing resources on demand. However, in current virtualized systems resource allocation is mostly static. Resources are allocated during VM instantiation and any change in workload leading to significant increase or decrease in resources is handled by VM migration. Hence, cloud users tend to characterize their workloads at a coarse grained level which potentially leads to under-utilized VM resources or under performing application. A more flexible and adaptive resource allocation mechanism would benefit variable workloads, such as those characterized by web servers. In this paper, we present an elastic resources framework for IaaS cloud layer that addresses this need. The framework provisions for application workload forecasting engine, that predicts at run-time the expected demand, which is input to the resource manager to modulate resource allocation based on the predicted demand. Based on the prediction errors, resources can be over-allocated or under-allocated as compared to the actual demand made by the application. Over-allocation leads to unused resources and under allocation could cause under performance. To strike a good trade-off between over-allocation and under-performance we derive an excess cost model. In this model excess resources allocated are captured as over-allocation cost and under-allocation is captured as a penalty cost for violating application service level agreement (SLA). Confidence interval for predicted workload is used to minimize this excess cost with minimal effect on SLA violations. An example case-study for an academic institute web server workload is presented. Using the confidence interval to minimize excess cost, we achieve significant reduction in resource allocation requirement while restricting application SLA violations to below 2-3%.
Mohit Dhingra, J. Lakshmi, S. K. Nandy 0001, Chiranjib Bhattacharyya, K. Gopinath
IEEE CLOUD5
2013 Subtle Topic Models and Discovering Subtly Manifested Software Concerns Automatically
abstract
In a recent pioneering approach LDA was used to discover cross cutting concerns(CCC) automatically from software codebases. LDA though successful in detecting prominent concerns, fails to detect many useful CCCs including ones that may be heavily executed but elude discovery because they do not have a strong prevalence in source-code. We pose this problem as that of discovering topics that rarely occur in individual documents, which we will refer to as subtle topics. Recently an interesting approach, namely focused topic models(FTM) was proposed for detecting rare topics. FTM, though successful in detecting topics which occur prominently in very few documents, is unable to detect subtle topics. Discovering subtle topics thus remains an important open problem. To address this issue we propose subtle topic models(STM). STM uses a generalized stick breaking process(GSBP) as a prior for defining multiple distributions over topics. This hierarchical structure on topics allows STM to discover rare topics beyond the capabilities of FTM. The associated inference is non-standard and is solved by exploiting the relationship between GSBP and generalized Dirichlet distribution. Empirical results show that STM is able to discover subtle CCC in two benchmark code-bases, a feat which is beyond the scope of existing topic models, thus demonstrating the potential of the model in automated concern discovery, a known difficult problem in Software Engineering. Furthermore it is observed that even in general text corpora STM outperforms the state of art in discovering subtle topics.
Mrinal Kanti Das, Suparna Bhattacharya, Chiranjib Bhattacharyya, K. Gopinath
ICML (2)4
2013 Combining concern input with program analysis for bloat detection
abstract
Framework based software tends to get bloated by accumulating optional features (or concerns) just-in-case they are needed. The good news is that such feature bloat need not always cause runtime execution bloat. The bad news is that often enough, only a few statements from an optional concern may cause execution bloat that may result in as much as 50% runtime overhead.
Suparna Bhattacharya, K. Gopinath, Mangala Gowri Nanda
OOPSLA2
2012 LoadIQ: Learning to Identify Workload Phases from a Live Storage Trace
Pankaj Pipada, Achintya Kundu, K. Gopinath, Chiranjib Bhattacharyya, Sai Susarla, P. C. Nagesh
HotStorage3
2012 Does lean imply green?: a study of the power performance implications of Java runtime bloat
abstract
The presence of software bloat in large flexible software systems can hurt energy efficiency. However, identifying and mitigating bloat is fairly effort intensive. To enable such efforts to be directed where there is a substantial potential for energy savings, we investigate the impact of bloat on power consumption under different situations. We conduct the first systematic experimental study of the joint power-performance implications of bloat across a range of hardware and software configurations on modern server platforms. The study employs controlled experiments to expose different effects of a common type of Java runtime bloat, excess temporary objects, in the context of the SPECPower_ssj2008 workload. We introduce the notion of equi-performance power reduction to characterize the impact, in addition to peak power comparisons. The results show a wide variation in energy savings from bloat reduction across these configurations. Energy efficiency benefits at peak performance tend to be most pronounced when bloat affects a performance bottleneck and non-bloated resources have low energy-proportionality. Equi-performance power savings are highest when bloated resources have a high degree of energy proportionality.
Suparna Bhattacharya, Karthick Rajamani, K. Gopinath
SIGMETRICS3
2011 Reuse, Recycle to De-bloat Software
Suparna Bhattacharya, Mangala Gowri Nanda, K. Gopinath
ECOOP3
2011 Virtually Cool Ternary Content Addressable Memory
Suparna Bhattacharya, K. Gopinath
HotOS2
2011 PRESIDIO: A Framework for Efficient Archival Data Storage
abstract
The ever-increasing volume of archival data that needs to be reliably retained for long periods of time and the decreasing costs of disk storage, memory, and processing have motivated the design of low-cost, high-efficiency disk-based storage systems. However, managed disk storage is still expensive. To further lower the cost, redundancy can be eliminated with the use of interfile and intrafile data compression. However, it is not clear what the optimal strategy for compressing data is, given the diverse collections of data. To create a scalable archival storage system that efficiently stores diverse data, we present PRESIDIO, a framework that selects from different space-reduction efficent storage methods (ESMs) to detect similarity and reduce or eliminate redundancy when storing objects. In addition, the framework uses a virtualized content addressable store (VCAS) that hides from the user the complexity of knowing which space-efficient techniques are used, including chunk-based deduplication or delta compression. Storing and retrieving objects are polymorphic operations independent of their content-based address. A new technique, harmonic super-fingerprinting, is also used for obtaining successively more accurate (but also more costly) measures of similarity to identify the existing objects in a very large data set that are most similar to an incoming new object. The PRESIDIO design, when reported earlier, had comprehensively introduced for the first time the notion of deduplication, which is now being offered as a service in storage systems by major vendors. As an aid to the design of such systems, we evaluate and present various parameters that affect the efficiency of a storage system using empirical data.
Lawrence You, Kristal T. Pollack, Darrell D. E. Long, K. Gopinath
ACM Trans. Storage4
2010 Discovery of Application Workloads from Network File Traces
Neeraja J. Yadwadkar, Chiranjib Bhattacharyya, K. Gopinath, Thirumale Niranjan, Sai Susarla
FAST3
2008 MAC Design for Heterogeneous Application Support in OFDM Based Wireless Systems
abstract
With the increasing adoption of wireless technology, it is reasonable to expect an increase in the demand for supporting both real-time multimedia and high rate reliable data services. Next generation wireless systems employ Orthogonal Frequency Division Multiplexing (OFDM) physical layer owing to the high data rate transmissions that are possible without increase in bandwidth. Towards improving the performance of these systems, we look at the design of resource allocation algorithms at medium-access iayer, and their impact on higher layers. While TCP-based elastic traffic needs reliable transport, UDP-based real-time applications have stringent delay and rate requirements. The MAC algorithms while catering to the heterogeneous service needs of these higher layers, tradeoff between maximizing the system capacity and providing fairness among users. The novelty of this work is the proposal of various channel-aware resource allocation algorithms at the MAC layer, which can result in significant performance gains in an OFDM based wireless system.
Prashanth L. A., K. Gopinath
CCNC3
2008 Structure and Interpretation of Computer Programs
abstract
Call graphs depict the static, caller-callee relation between "functions " in a program. With most source/target languages supporting functions as the primitive unit of composition, call graphs naturally form the fundamental control flow representation available to understand/develop software. They are also the substrate on which various inter- procedural analyses are performed and are integral part of program comprehension/testing. Given their universality and usefulness, it is imperative to ask if call graphs exhibit any intrinsic graph theoretic features - across versions, program domains and source languages. This work is an attempt to answer these questions: we present and investigate a set of meaningful graph measures that help us understand call graphs better; we establish how these measures correlate, if any, across different languages and program domains; we also assess the overall, language independent software quality by suitably interpreting these measures.
Ganesh M. Narayan, K. Gopinath, Sridhar Varadarajan
TASE2
2007 Optimizing multimedia experience in a thin client environment for a resource constrained processor
abstract
In this paper, we study how TCP and UDP flows interact with each other when the end system is a CPU resource constrained thin client. The problem addressed is twofold, 1) the throughput of TCP flows degrades severely in the presence of heavily loaded UDP flows 2) fairness and minimum QoS requirements of UDP are not maintained. First, we identify the factors affecting the TCP throughput by providing an in-depth analysis of end to end delay and packet loss variations. The results obtained from the first part leads us to our second contribution. We propose and study the use of an algorithm that ensures fairness across flows. The algorithm improves the performance of TCP flows in the presence of multiple UDP flows admitted under an admission algorithm and maintains the minimum QoS requirements of the UDP flows. The advantage of the algorithm is that it requires no changes to TCP/IP stack and control is achieved through receiver window control.
G. A. Ramanujan, Amit Thawani, K. Gopinath
IWCMC4
2007 Performance Evaluation of Multiple TCP connections in iSCSI
Bhargava Kumar K, Ganesh M. Narayan, K. Gopinath
MSST3
2007 Recovery from DoS Attacks in MIPv6: Modeling and Validation
abstract
Denial-of-Service (DoS) attacks form a very important category of security threats that are prevalent in MIPv6 (Mobile Internet Protocol version 6) today. Many schemes have been proposed to alleviate such threats, including one of our own [9]. However, reasoning about the correctness of such protocols is not trivial. In addition, new solutions to mitigate attacks may need to be deployed in the network on a frequent basis as and when attacks are detected, as it is practically impossible to anticipate all attacks and provide solutions in advance. This makes it necessary to validate the solutions in a timely manner before deployment in the real network. However, threshold schemes needed in group protocols make analysis complex. Model checking threshold-based group protocols that employ cryptography have not been successful so far. Here, we propose a new simulation based approach for validation using a tool called FRAMOGR that supports executable specification of group protocols that use cryptography. FRAMOGR allows one to specify attackers and track probability distributions of values or paths. We believe that infrastructure such as FRAMOGR would be required in future for validating new group based threshold protocols that may be needed for making MIPv6 more robust.
Manish C. Kumar, K. Gopinath
SEFM2
2006 An Extended Verifiable Secret Redistribution Protocol for Archival Systems
abstract
Existing protocols for archival systems make use of verifiability of shares in conjunction with a proactive secret sharing scheme to achieve high availability and long term confidentiality, besides data integrity. In this paper, we extend an existing protocol (Wong et al. [2002]) to take care of more realistic situations. For example, it is assumed in the protocol of Wong et al. that the recipients of the secret shares are all trustworthy; we relax this by requiring that only a majority is trustworthy.
V. H. Gupta, K. Gopinath
ARES2
2006 Proactive Leader Election in Asynchronous Shared Memory Systems
M. C. Dharmadeep, K. Gopinath
ATVA2
2005 Improved Probabilistic Models for 802.11 Protocol Verification
Amitabha Roy 0002, K. Gopinath
CAV2
2005 Evaluation of Advanced TCP Stacks in the iSCSI Environment using Simulation Model
abstract
Enterprise storage demands have overwhelmed traditional storage mechanisms and have led to the development of storage area networks (SANs). This has resulted in the design of SCSI transport protocols that encapsulate SCSI commands and data for transfer over the network. Fiber channel protocol was the first such protocol that used gigabit per second speed links to carry SCSI commands and data over long distances. However, with the emergence of gigabit Ethernet and the iSCSI (Internet SCSI) protocol that maps the SCSI block oriented storage data over TCP/IP and enables storage devices to be accessed over standard Ethernet based TCP/IP networks, reduction in costs and a unified network infrastructure can be achieved. The iSCSI data flow is regulated by the TCP congestion control algorithm. The standard TCP Reno congestion control algorithm substantially under utilizes the network bandwidth over high speed connections for most applications. To address this limitation of TCP, variants of the TCP congestion control algorithm, designed for high bandwidth networks, such as FAST TCP, BIC-TCP, H-TCP, and Scalable TCP have been proposed. We use simulations and experiments to compare the performance of these TCP variants in an iSCSI environment. Our results indicate that H-TCP outperforms the other TCP variants in an iSCSI environment. However H-TCP results in unfair sharing of network bandwidth when simultaneous multiple flows exist. Performance obtained using BIC-TCP is only second to that using H-TCP and it is relatively fairer as compared to H-TCP.
Girish Motwani, K. Gopinath
MSST2
2004 iSAN - An Intelligent Storage Area Network Architecture
Ganesh M. Narayan, K. Gopinath
HiPC2
2001 EASN: Integrating ASN.1 and Model Checking
Vivek K. Shanbhag, K. Gopinath, Markku Turunen, Ari Ahtiainen, Matti Luukkainen
CAV2
2001 Verification of a Leader Election Algorithm in Timed Asynchronous Systems
Neeraj Jaggi, K. Gopinath
FSTTCS2
2000 A Multi-Tier RAID Storage System with RAID1 and RAID5
abstract
Redundant Arrays of Inexpensive Disks (RAID) is a popular technique used to improve the reliability and performance of secondary storage. Of various levels of RAID discussed, RAID1 and RAID5 have become more popular. Mirroring or RAID1 maintains multiple copies of the data, generally provides best performance and is easier to configure. Rotating parity scheme or RAID5 is the least expensive RAID scheme with good large update performance. It suffers from poor small update performance and performance drops sharply when a diskfails and the array enters degraded mode. Configuring RAID5 is more involved. This paper presents the design and implementation of a host-based driver for a multi-tier RAID storage system, currently with 2 tiers: a small RAID1 tier and a larger RAID5 tier. Based on access patterns, the driver automatically migrates frequently accessed data to RAID1 while demoting not so frequently accessed data to RAID5. The prototype provides reliable persistence semantics for data migration between the tiers using ordered updates. Mechanisms are separated from policies through an API so that any desired policy can be implemented in trusted user processes. Finally, we present comparison of the performance of our system with comparable systems using striping and RAID5.
Nitin Muppalaneni, K. Gopinath
IPDPS2
1999 Combining Conditional Constant Propagation and Interprocedural Alias Analysis
K. Gopinath, K. S. Nandakumar
HiPC1
1998 Formal Verification of an O. S. Submodule
N. S. Pendharkar, K. Gopinath
FSTTCS2
1998 Data structure distribution and multi-threading of Linux file system for multiprocessors
abstract
The standard Linux design assumes a uniprocessor architecture. Allowing several processors to execute simultaneously in the kernel mode on behalf of different processes can cause consistency problems unless appropriate exclusion mechanisms are used. In addition, if the file system data structures are not distributed, performance can be affected. We discuss a multiprocessor file system design for Linux ext2fs with various data structures, such as super block, inodes, buffer cache, directory cache (name cache), distributed with respect to different processors with appropriate exclusion mechanisms.
Anish Sheth, K. Gopinath
HiPC2
1997 Characterizing vulnerability of parallelism to resource constraints
abstract
The theoretical available instruction level parallelism in most benchmark is very high. Vulnerability is related to the difficulty with which we can extract this parallelism with finite resources. This study characterizes the vulnerability of parallelism to resource constraints by scheduling dynamic dependence graphs (DDGs) from traces of several benchmarks using different scheduling algorithms and different number of functional units. It is observed that the execution time of the DDGs does not vary significantly with low-level scheduling algorithms like lazy, slack, etc. Measures of vulnerability based on slack and load were also considered. Although Accslk-Load, which uses a combination of accurate slack and load to make a prediction, has a prediction accuracy of about 85%, the prediction rate is only 42%. On the other hand, even though the prediction accuracy of /spl sigma/(L/sup x/), the standard deviation in the load, is not as high, there is a prediction in all the cases. The DDG execution time is also found to be most vulnerable to the functional unit with the greatest /spl sigma/(L/sub x/).
V. Vivekanand, K. Gopinath, Pradeep Dubey
HiPC2
1997 A C++ Simulator Generator from Graphical Specifications
abstract
Many languages for computer systems simulation (like GPSS and CSim) use a stochastic model of systems with the provision of adding procedural code for those aspects of the system that cannot be captured easily by a stochastic model. However, they do not support the hierachical simulation of complex systems well. Complex computer systems may have to be simulated at various levels of abstraction in the interests of tractability: the flexibility of being able to freely move between the different levels of abstraction is very desirable. For example, in the area of computer architecture, one might have analytical models, detailed simulation models and trace-driven models. In addition, these languages do not have user-friendly interfaces for specification of the simulated system. In this paper, we discuss the design and implementation of a simulation package for hierachical simulation of non-real-time computer systems: a Simulator Generator from a Graphical System Specification (SIGGSYS}). A new language for system specification has been designed. In addition, the package has the following components: • A graphical user interface to aid specification of the system to be simulated. • A rear end that generates C++ code that implements a simulator for the specified system. • A complete object library along with the header files that implement a functionally complete set of C++ base classes which can be built upon. C++ has been chosen as the intermediate language so that the modeller can use its support for object oriented programming. © 1997 John Wiley & Sons, Ltd.
Vivek K. Shanbhag, K. Gopinath
Softw. Pract. Exp.2
1996 Program analysis for page size selection
abstract
To support high performance architectures with multiple page sizes, it is necessary to assign proper page sizes for array memory in order to improve TLB performance as well as reduce memory contention during program execution. Typically, while a smaller page size causes higher TLB contention, a larger page size causes higher memory contention and fragmentation but also has the effect of prefetching pages required in future thereby reducing the number of cold page faults. Each array in a program contributes to these costs/benefits depending upon how it is referenced in the program. The page size assignment analysis determines a proper page size for every array by analyzing memory reference patterns (which is shown to be NP-hard). We discuss various policies that can be followed for page size assignment in order to maximize performance along with cost models and present algorithms for page size selection.
K. Gopinath, Aniruddha P. Bhutkar
HiPC1
1994 Performance of Switch Blocking on Multithreaded Architectures
abstract
Block multithreaded architectures tolerate large memory and synchronization latencies by switching contexts on every remote-memory-access or on a failed synchronization request. We study the performance of a waiting mechanism called switch-blocking where waiting threads are disabled (but not unloaded) and signalled at the completion of the wait in comparison with switch-spinning where waiting threads poll and execute in a round-robin fashion. We present an implementation of switch-blocking on a simulator for Alewife (a block multithreaded machine) for both remote memory accesses and synchronization operations and discuss results from the simulator. Our results indicate that switch-blocking has the same problems that switch-spinning has under heavy lock contention and that support for switch-blocking for remote memory accesses may not be judicious at current range of memory access times but may be so in the future due to its strong interactions with synchronization operations.
K. Gopinath, M. K. Krishna Narasimhan, Beng-Hong Lim, Anant Agarwal
ICPP (1)1
1991 Memory Models Compiler Optimizations and P-RISC
K. Gopinath
ICPP (1)1
1989 Copy Elimination in Functional Languages
abstract
Copy elimination is an important optimization for compiling functional languages. Copies arise because these languages lack the concepts of state and variable; hence updating an object involves a copy in a naive implementation. Copies are also possible if proper targeting has not been carried out inside functions and across function calls. Targeting is the proper selection of a storage area for evaluating an expression. By abstracting a collection of functions by a target operator, we compute targets of function bodies that can then be used to define an optimized interpreter to eliminate copies due to updates and copies across function calls. The language we consider is typed lambda calculus with higher-order functions and special constructs for array operations. Our approach can eliminate copies in divide and conquer problems like quicksort and bitonic sort that previous approaches could not handle.
K. Gopinath, John L. Hennessy
POPL1