Chandramohan A. Thekkath

dblp:13/3927 · DBLP profile ↗
← Back
27ranked-venue papers
7as first author
3since 2021 · last 2024
0009-0004-9924-2428ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 12 · 3 first-author · 3 since 2021Software engineering, systems software and programming languages · 12 · 4 first-authorComputer networks · 6 · 2 first-authorDatabases, data management, data science and information retrieval · 2Security and privacy · 1
YearPublicationVenuePosition
2024 Pecan: Cost-Efficient ML Data Preprocessing with Automatic Transformation Ordering and Hybrid Placement
Dan Graur, Oto Mraz, Muyu Li, Mohammad Sepehr Pourghannad, Chandramohan A. Thekkath, Ana Klimovic
USENIX ATC5
2023 tf.data service: A Case for Disaggregating ML Input Data Processing
abstract
Machine learning (ML) computations commonly execute on expensive specialized hardware, such as GPUs and TPUs, which provide high FLOPs and performance-per-watt. For cost efficiency, it is essential to keep these accelerators highly utilized. This requires preprocessing input data at the rate at which the accelerators can ingest and perform ML computations on the data. To avoid data stalls, the host CPU and RAM required for input data processing per accelerator core used for ML computations varies across jobs. Hence, the traditional approach of processing input data on ML accelerator hosts with a fixed hardware ratio leads to either under-utilizing the accelerators or the host CPU and RAM. In this paper, we address these concerns by building a disaggregated ML data processing system.
Andrew Audibert, Yang Chen 0003, Dan Graur, Ana Klimovic, Jiri Simsa, Chandramohan A. Thekkath
SoCC6
2022 Cachew: Machine Learning Input Data Processing as a Service
Dan Graur, Damien Aymon, Dan Kluser, Tanguy Albrici, Chandramohan A. Thekkath, Ana Klimovic
USENIX ATC5
2011 Chimera: data sharing flexibility, shared nothing simplicity
abstract
The current database market is fairly evenly split between shared nothing and data sharing systems. While shared nothing systems are easier to build and scale, data sharing systems have advantages in load balancing. In this paper we explore adding data sharing functionality as an extension to a shared nothing database system. Our approach isolates the data sharing functionality from the rest of the system and relies on well-studied, robust techniques to provide the data sharing extension. This reduces the difficulty in providing data sharing functionality, yet provides much of the flexibility of a data sharing system. We present the design and implementation of Chimera -- a hybrid database system, targeted at load balancing for many workloads, and scale-out for read-mostly workloads. The results of our experiments demonstrate that we can achieve almost linear scalability and effective load balancing with less than 2% overhead during normal operation.
Umar Farooq Minhas, David B. Lomet, Chandramohan A. Thekkath
IDEAS3
2010 Nectar: Automatic Management of Data and Computation in Datacenters
Pradeep Kumar Gunda, Lenin Ravindranath, Chandramohan A. Thekkath, Li Zhuang
OSDI3
2010 StarTrack Next Generation: A Scalable Infrastructure for Track-Based Applications
Maya Haridasan, Iqbal Mohomed, Douglas B. Terry, Chandramohan A. Thekkath, Li Zhang 0001
OSDI4
2009 StarTrack: a framework for enabling track-based applications
abstract
Mobile devices are increasingly equipped with hardware and software services allowing them to determine their locations, but support for building location-aware applications remains rudimentary. This paper proposes tracks of location coordinates as a high-level abstraction for a new class of mobile applications including ride sharing, location-based collaboration, and health monitoring. Each track is a sequence of entries recording a person's time, location, and application-specific data. StarTrack provides applications with a comprehensive set of operations for recording, comparing, clustering and querying tracks. StarTrack can efficiently operate on thousands of tracks.
Ganesh Ananthanarayanan, Maya Haridasan, Iqbal Mohomed, Douglas B. Terry, Chandramohan A. Thekkath
MobiSys5
2008 Niobe: A practical replication protocol
abstract
The task of consistently and reliably replicating data is fundamental in distributed systems, and numerous existing protocols are able to achieve such replication efficiently. When called on to build a large-scale enterprise storage system with built-in replication, we were therefore surprised to discover that no existing protocols met our requirements. As a result, we designed and deployed a new replication protocol called Niobe . Niobe is in the primary-backup family of protocols, and shares many similarities with other protocols in this family. But we believe Niobe is significantly more practical for large-scale enterprise storage than previously published protocols. In particular, Niobe is simple, flexible, has rigorously proven yet simply stated consistency guarantees, and exhibits excellent performance. Niobe has been deployed as the backend for a commercial Internet service; its consistency properties have been proved formally from first principles, and further verified using the TLA + specification language. We describe the protocol itself, the system built to deploy it, and some of our experiences in doing so.
John MacCormick, Chandramohan A. Thekkath, Marcus Jager, Kristof Roomp, Lidong Zhou, Ryan S. Peterson
ACM Trans. Storage2
2007 COMBINE: leveraging the power of wireless peers through collaborative downloading
abstract
Mobile devices are increasingly equipped with multiple network interfaces: Wireless Local Area Network (WLAN) interfaces for local connectivity and Wireless Wide Area Network (WWAN) interfaces for wide-area connectivity. The WWAN typically provides much wider coverage but much lower speeds than the WLAN. To address this dichotomy, we present COMBINE, a system for collaborative downloading wherein devices that are within WLAN range pool together their WWAN links, significantly increasing the effective speed available to them.
Ganesh Ananthanarayanan, Venkat N. Padmanabhan, Lenin Ravindranath, Chandramohan A. Thekkath
MobiSys4
2007 Graceful degradation via versions: specifications and implementations
abstract
Correctness of a fault-tolerant system hinges on the failure model, which typically constrains the number of concurrent failures in the system. These assumptions are sometimes violated in practice, inevitably leading to degraded system behavior that deviates from the system's specification and even causing complete unavailability of the system.
Lidong Zhou, Vijayan Prabhakaran, Venugopalan Ramasubramanian, Roy Levin, Chandramohan A. Thekkath
PODC5
2005 SenSlide: a sensor network based landslide prediction aystem
Anmol Sheth, Kalyan Tejaswi, Prakshep Mehta, Chandresh Parekh, Rajul Bansal, S. N. Merchant, T. N. Singh 0001, Uday B. Desai, Chandramohan A. Thekkath, K. Toyama
SenSys9
2004 Boxwood: Abstractions as the Foundation for Storage Infrastructure
John MacCormick, Nick Murphy, Marc Najork, Chandramohan A. Thekkath, Lidong Zhou
OSDI4
2003 Block-Level Security for Network-Attached Disks
Marcos K. Aguilera, Minwen Ji, Mark Lillibridge, John MacCormick, Erwin Oertli, David G. Andersen, Michael Burrows, Timothy P. Mann, Chandramohan A. Thekkath
FAST9
2003 Implementing an untrusted operating system on trusted hardware
abstract
Recently, there has been considerable interest in providing "trusted computing platforms" using hardware~---~TCPA and Palladium being the most publicly visible examples. In this paper we discuss our experience with building such a platform using a traditional time-sharing operating system executing on XOM~---~a processor architecture that provides copy protection and tamper-resistance functions. In XOM, only the processor is trusted; main memory and the operating system are not trusted.Our operating system (XOMOS) manages hardware resources for applications that don't trust it. This requires a division of responsibilities between the operating system and hardware that is unlike previous systems. We describe techniques for providing traditional operating systems services in this context.Since an implementation of a XOM processor does not exist, we use SimOS to simulate the hardware. We modify IRIX 6.5, a commercially available operating system to create xomos. We are then able to analyze the performance and implementation overheads of running an untrusted operating system on trusted hardware.
David Lie, Chandramohan A. Thekkath, Mark Horowitz
SOSP2
2003 Specifying and Verifying Hardware for Tamper-Resistant Software
abstract
We specify a hardware architecture that supports tamper-resistant software by identifying an "idealized" model, which gives the abstracted actions available to a single user program. This idealized model is compared to a concrete "actual" model that includes actions of an adversarial operating system. The architecture is verified by using a finite-state enumeration tool (a model checker) to compare executions of the idealized and actual models. In this approach, software tampering occurs if the system can enter a state where one model is inconsistent with the other in performing the verification, we detected a replay attack scenario and were able to verify the security of our solution to the problem. Our methods were also able to verify that all actions in the architecture are required, as well as come up with a set of constraints on the operating system to guarantee liveness for users.
David Lie, John C. Mitchell, Chandramohan A. Thekkath, Mark Horowitz
S&P3
2000 Architectural Support for Copy and Tamper Resistant Software
abstract
Although there have been attempts to develop code transformations that yield tamper-resistant software, no reliable software-only methods are know. This paper studies the hardware implementation of a form of execute-only memory (XOM) that allows instructions stored in memory to be executed but not otherwise manipulated. To support XOM code we use a machine that supports internal compartments---a process in one compartment cannot read data from another compartment. All data that leaves the machine is encrypted, since we assume external memory is not secure. The design of this machine poses some interesting trade-offs between security, efficiency, and flexibility. We explore some of the potential security issues as one pushes the machine to become more efficient and flexible. Although security carries a performance penalty, our analysis indicates that it is possible to create a normal multi-tasking machine where nearly all applications can be run in XOM mode. While a virtual XOM machine is possible, the underlying hardware needs to support a unique private key, private memory, and traps on cache misses. For efficient operation, hardware assist to provide fast symmetric ciphers is also required.
David Lie, Chandramohan A. Thekkath, Mark Mitchell, Patrick Lincoln, Dan Boneh, John C. Mitchell, Mark Horowitz
ASPLOS2
2000 SmartBridge: A scalable bridge architecture
abstract
As the number of hosts attached to a network increases beyond what can be connected by a single local area network (LAN), forwarding packets between hosts on different LANs becomes an issue. Two common solutions to the forwarding problem are IP routing and spanning tree bridging. IP routing scales well, but imposes the administrative burden of managing subnets and assigning addresses. Spanning tree bridging, in contrast, requires no administration, but often does not perform well in a large network, because too much traffic must detour toward the root of the spanning tree, wasting link bandwidth.
Thomas L. Rodeheffer, Chandramohan A. Thekkath, Darrell C. Anderson
SIGCOMM2
1997 Frangipani: A Scalable Distributed File System
abstract
article Frangipani: a scalable distributed file system Share on Authors: Chandramohan A. Thekkath Systems Research Center, Digital Equipment Corporation, 130 Lytton Ave, Palo Alto, CA Systems Research Center, Digital Equipment Corporation, 130 Lytton Ave, Palo Alto, CAView Profile , Timothy Mann Systems Research Center, Digital Equipment Corporation, 130 Lytton Ave, Palo Alto, CA Systems Research Center, Digital Equipment Corporation, 130 Lytton Ave, Palo Alto, CAView Profile , Edward K. Lee Systems Research Center, Digital Equipment Corporation, 130 Lytton Ave, Palo Alto, CA Systems Research Center, Digital Equipment Corporation, 130 Lytton Ave, Palo Alto, CAView Profile Authors Info & Claims ACM SIGOPS Operating Systems ReviewVolume 31Issue 5Dec. 1997 pp 224–237https://doi.org/10.1145/269005.266694Online:01 October 1997Publication History 250citation3,355DownloadsMetricsTotal Citations250Total Downloads3,355Last 12 Months141Last 6 weeks25 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access
Chandramohan A. Thekkath, Timothy P. Mann, Edward K. Lee 0001
SOSP1
1996 Petal: Distributed Virtual Disks
abstract
The ideal storage system is globally accessible, always available, provides unlimited performance and capacity for a large number of clients, and requires no management. This paper describes the design, implementation, and performance of Petal, a system that attempts to approximate this ideal in practice through a novel combination of features. Petal consists of a collection of network-connected servers that cooperatively manage a pool of physical disks. To a Petal client, this collection appears as a highly available block-level storage system that provides large abstract containers called virtual disks. A virtual disk is globally accessible to all Petal clients on the network. A client can create a virtual disk on demand to tap the entire capacity and performance of the underlying physical resources. Furthermore, additional resources, such as servers and disks, can be automatically incorporated into Petal.We have an initial Petal prototype consisting of four 225 MHz DEC 3000/700 workstations running Digital Unix and connected by a 155 Mbit/s ATM network. The prototype provides clients with virtual disks that tolerate and recover from disk, server, and network failures. Latency is comparable to a locally attached disk, and throughput scales with the number of servers. The prototype can achieve I/O rates of up to 3150 requests/sec and bandwidth up to 43.1 Mbytes/sec.
Edward K. Lee 0001, Chandramohan A. Thekkath
ASPLOS2
1996 Shasta: A Low Overhead, Software-Only Approach for Supporting Fine-Grain Shared Memory
abstract
This paper describes Shasta, a system that supports a shared address space in software on clusters of computers with physically distributed memory. A unique aspect of Shasta compared to most other software distributed shared memory systems is that shared data can be kept coherent at a fine granularity. In addition, the system allows the coherence granularity to vary across different shared data structures in a single application. Shasta implements the shared address space by transparently rewriting the application executable to intercept loads and stores. For each shared load or store, the inserted code checks to see if the data is available locally and communicates with other processors if necessary. The system uses numerous techniques to reduce the run-time overhead of these checks. Since Shasta is implemented entirely in software, it also provides tremendous flexibility in supporting different types of cache coherence protocols. We have implemented an efficient cache coherence protocol that incorporates a number of optimizations, including support for multiple communication granularities and use of relaxed memory models. This system is fully functional and runs on a cluster of Alpha workstations.The primary focus of this paper is to describe the techniques used in Shasta to reduce the checking overhead for supporting fine granularity sharing in software. These techniques include careful layout of the shared address space, scheduling the checking code for efficient execution on modern processors, using a simple method that checks loads using only the value loaded, reducing the extra cache misses caused by the checking code, and combining the checks for multiple loads and stores. To characterize the effect of these techniques, we present detailed performance results for the SPLASH-2 applications running on an Alpha processor. Without our optimizations, the checking overheads are excessively high, exceeding 100% for several applications. However, our techniques are effective in reducing these overheads to a range of 5% to 35% for almost all of the applications. We also describe our coherence protocol and present some preliminary results on the parallel performance of several applications running on our workstation cluster. Our experience so far indicates that once the cost of checking memory accesses is reduced using our techniques, the Shasta approach is an attractive software solution for supporting a shared address space with fine-grain access to data.
Daniel J. Scales, Kourosh Gharachorloo, Chandramohan A. Thekkath
ASPLOS3
1995 Implementing Global Memory Management in a Workstation Cluster
abstract
Advances in network and processor technology have greatly changed the communication and computational power of local-area workstation clusters.However, operating systems still treat workstation clusters as a collection of loosely-connected processors,where each workstation acts as an autonomous and independent agent.This operating system structure makes it difficult to exploit the characteristics of current clusters, such as low-latency communication, huge primary memories, and high-speed processors, in order to improve the performance of cluster applications.This paper describes the design and implementation of global memory management in a workstation cluster.Our objective is to use a single, unified, but distributed memory management algorithm at the lowest level of the operating system.By managing memory globally at this level, all system-and higher-level software, including VM, file systems, transaction systems, and user applications, can benefit from available cluster memory.We have implemented our algorithm in the OSF/1 operating system running on an ATM-connected cluster of DEC Alpha workstations.Our measurements show that on a suite of memory-intensive programs, our system improves performance by a factor of 1.5 to 3.5.We also show that our algorithm has a performance advantage over others that have been proposed in the past.
Michael J. Feeley, William E. Morgan, Frédéric H. Pighin, Anna R. Karlin, Henry M. Levy, Chandramohan A. Thekkath
SOSP6
1994 Hardware and Software Support for Efficient Exception Handling
abstract
Program-synchronous exceptions, for example, breakpoints, watchpoints, illegal opcodes, and memory access violations, provide information about exceptional conditions, interrupting the program and vectoring to an operating system handler. Over the last decade, however, programs and run-time systems have increasingly employed these mechanisms as a performance optimization to detect normal and expected conditions. Unfortunately, current architecture and operating system structures are designed for exceptional or erroneous conditions, where performance is of secondary importance, rather than normal conditions. Consequently, this has limited the practicality of such hardware-based detection mechanisms.
Chandramohan A. Thekkath, Henry M. Levy
ASPLOS1
1994 Separating Data and Control Transfer in Distributed Operating Systems
abstract
Advances in processor architecture and technology have resulted in workstations in the 100+ MIPS range. As well, newer local-area networks such as ATM promise a ten- to hundred-fold increase in throughput, much reduced latency, greater scalability, and greatly increased reliability, when compared to current LANs such as Ethernet.
Chandramohan A. Thekkath, Henry M. Levy, Edward D. Lazowska
ASPLOS1
1994 Techniques for File System Simulation
abstract
Abstract Careful simulation‐based evaluation plays an important role in the design of file and disk systems. We describe here a particular approach to such evaluations that combines techniques in workload synthesis, file system modeling, and detailed disk behavior modeling. Together, these make feasible the detailed simulation of I/O hardware and file system software. In particular, using the techniques described here is likely to make comparative file system studies more accurate. In addition to these specific contributions, the paper makes two broader points. First, it argues that detailed models are appropriate and necessary in many cases. Second, it demonstrates that detailed models need not be difficult or time consuming to construct or execute.
Chandramohan A. Thekkath, John Wilkes, Edward D. Lazowska
Softw. Pract. Exp.1
1993 Implementing Network Protocols at User Level
abstract
Traditionally, network software has been structured in a monolithic fashion with all protocol stacks executing either within the kernel or in a single trusted user-level server. This organization is motivated by performance and security concerns. However, considerations of code maintenance, ease of debugging, customization, and the simultaneous existence of multiple protocols argue for separating the implementations into more manageable user-level libraries of protocols. This paper describes the design and implementation of transport protocols as user-level libraries.We begin by motivating the need for protocol implementations as user-level libraries and placing our approach in the context of previous work. We then describe our alternative to monolithic protocol organization, which has been implemented on Mach workstations connected not only to traditional Ethernet, but also to a more modern network, the DEC SRC ANI. Based on our experience, we discuss the implications for host-network interface design and for overall system structure to support efficient user-level implementations of network protocols.
Chandramohan A. Thekkath, Thu D. Nguyen, Evelyn Moy, Edward D. Lazowska
SIGCOMM1
1993 Limits to Low-Latency Communication on High-Speed Networks
abstract
The throughput of local area networks is rapidly increasing. For example, the bandwidth of new ATM networks and FDDI token rings is an order of magnitude greater than that of Ethernets. Other network technologies promise a bandwidth increase of yet another order of magnitude in several years. However, in distributed systems, lowered latency rather than increased throughput is often of primary concern. This paper examines the system-level effects of newer high-speed network technologies on low-latency, cross-machine communications. To evaluate a number of influences, both hardware and software, we designed and implemented a new remote procedure call system targeted at providing low latency. We then ported this system to several hardware platforms (DECstation and SPARCstation) with several different networks and controllers (ATM, FDDI, and Ethernet). Comparing these systems allows us to explore the performance impact of alternative designs in the communication system with respect to achieving low latency, e.g., the network, the network controller, the hose architecture and cache system, and the kernel and user-level runtime software. Our RPC system, which achieves substantially reduced call times (170 μseconds on an ATM network using DECstation 5000/200 hosts), allows us to isolate those components of next-generation networks and controllers that still stand in the way of low-latency communication. We demonstrate that new-generation processor technology and software design can reduce small-packet RPC times to near network-imposed limits, making network and controller design more crucial than ever to achieving truly low-latency communication.
Chandramohan A. Thekkath, Henry M. Levy
ACM Trans. Comput. Syst.1
1993 Implementing network protocols at user level
abstract
Traditionally, network software has been structured in a monolithic fashion with all protocol stacks executing either within the kernel or in a single trusted user-level server. This organization is motivated by performance and security concerns. However, considerations of code maintenance, ease of debugging, customization, and the simultaneous existence of multiple protocols argue for separating the implementations into more manageable user-level libraries of protocols. The present paper describes the design and implementation of transport protocols as user-level libraries. The authors begin by motivating the need for protocol implementations as user-level libraries and placing their approach in the context of previous work. They then describe their alternative to monolithic protocol organization, which has been implemented on Mach workstations connected not only to traditional Ethernet, but also to a more modern network, the DEC SRC AN1. Based on the authors' experience, they discuss the implications for host-network interface design and for overall system structure to support efficient user-level implementations of network protocols.>
Chandramohan A. Thekkath, Thu D. Nguyen, Evelyn Moy, Edward D. Lazowska
IEEE/ACM Trans. Netw.1