Nishkam Ravi

dblp:r/NishkamRavi · DBLP profile ↗
← Back
14ranked-venue papers
7as first author
1since 2021 · last 2021
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 2 first-authorHuman-computer interaction and ubiquitous computing · 5 · 4 first-authorArtificial intelligence and machine learning · 3 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-authorSoftware engineering, systems software and programming languages · 1Databases, data management, data science and information retrieval · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
3 papers
Cloud and datacenter computing · 51% Parallel and multicore computing · 22% Hardware accelerators and domain-specific architectures · 17%
Databases, data mining, and information retrieval
1 paper
Query processing and optimization · 100%
Software engineering, system software, and programming languages
2 papers
Compilers and program optimization · 100%
Human-computer interaction and pervasive computing
3 papers
Ubiquitous computing and smart environments · 71% Interaction techniques and input · 22% User interface design and tools · 7%
Computer networks
3 papers
Internet of things and sensor networks · 28% Wireless networking · 28% Network management and operations · 24%

Topics — the 16 heaviest of 20, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Query processing and optimization
query optimization
0.512021
SparkCruise: Workload Optimization in Managed Spark Clusters at Microsoft · Proc. VLDB Endow. 2021
Cloud and datacenter computing
cluster resource management and scheduling
0.512021
SparkCruise: Workload Optimization in Managed Spark Clusters at Microsoft · Proc. VLDB Endow. 2021
Hardware accelerators and domain-specific architectures
many-core accelerator
0.212013
Semi-automatic restructuring of offloadable tasks for many-core accelerators · SC 2013
Parallel and multicore computing
work distribution
0.212013
Semi-automatic restructuring of offloadable tasks for many-core accelerators · SC 2013
Query processing and optimization › result reuse
computation reuse
0.112021
SparkCruise: Workload Optimization in Managed Spark Clusters at Microsoft · Proc. VLDB Endow. 2021
Ubiquitous computing and smart environments
context-aware computing
0.112008
Context-aware Battery Management for Mobile Phones · PerCom 2008
Interaction techniques and input
cross-device interaction
0.112006
Inverted Browser: A Novel Approach towards Display Symbiosis · PerCom 2006
Ubiquitous computing and smart environments › public displays
public display interaction
0.112006
Inverted Browser: A Novel Approach towards Display Symbiosis · PerCom 2006
GPUs and heterogeneous computing › GPU communication
host-device data transfer
0.112014
COMP: Compiler Optimizations for Manycore Processors · MICRO 2014
Ubiquitous computing and smart environments › context recognition
activity recognition
0.112005
Activity Recognition from Accelerometer Data · AAAI 2005
Internet of things and sensor networks
service discovery
0.112005
Accessing Ubiquitous Services Using Smart Phones · PerCom 2005
Wireless networking
service discovery protocol
0.112005
Accessing Ubiquitous Services Using Smart Phones · PerCom 2005
High-performance computing › performance optimization
auto-tuning
0.012013
Semi-automatic restructuring of offloadable tasks for many-core accelerators · SC 2013
Parallel and multicore computing › parallel scheduling
runtime scheduling
0.012013
Semi-automatic restructuring of offloadable tasks for many-core accelerators · SC 2013
User interface design and tools › interactive systems
web browser
0.012006
Inverted Browser: A Novel Approach towards Display Symbiosis · PerCom 2006
Privacy and data protection › anonymity
anonymous payment
0.012005
Accessing Ubiquitous Services Using Smart Phones · PerCom 2005

Methods — techniques the papers use, named apart from their topics

workload analysis · 1.0source-to-source compilation · 0.4compiler directives · 0.3location trace analysis · 0.2call log analysis · 0.2millicent scrips · 0.1instance-based state representation · 0.1web services · 0.1prototype implementation · 0.1accelerometer · 0.1
YearPublicationVenuePosition
2021 SparkCruise: Workload Optimization in Managed Spark Clusters at Microsoft
abstract
Today cloud companies offer fully managed Spark services. This has made it easy to onboard new customers but has also increased the volume of users and their workload sizes. However, both cloud providers and users lack the tools and time to optimize these massive workloads. To solve this problem, we designed SparkCruise that can help understand and optimize workload instances by adding a workload-driven feedback loop to the Spark query optimizer. In this paper, we present our approach to collecting and representing Spark query workloads and use it to improve the overall performance on the workload, all without requiring any access to user data. These methods scale with the number of workloads and apply learned feedback in an online fashion. We explain one specific workload optimization developed for computation reuse. We also share the detailed analysis of production Spark workloads and contrast them with the corresponding analysis of TPC-DS benchmark. To the best of our knowledge, this is the first study to share the analysis of large-scale production Spark SQL workloads.
Abhishek Roy 0008, Alekh Jindal, Priyanka Gomatam, Xiating Ouyang, Ashit Gosalia, Nishkam Ravi, Swinky Mann, Prakhar Jain
Proc. VLDB Endow.6
2014 Automating and optimizing data transfers for many-core coprocessors
abstract
Orchestrating data transfers between CPUs and a coprocessor manually is cumbersome, particularly for multi-dimensional arrays and other data structures with multi-level pointers, which are common in scientific computations. This work describes a system that includes both compile-time and runtime solutions for this problem, with the overarching goal of improving programmer productivity while maintaining performance.
Bin Ren 0002, Nishkam Ravi, Yi Yang 0018, Min Feng 0001, Gagan Agrawal, Srimat T. Chakradhar
ICS2
2014 COMP: Compiler Optimizations for Manycore Processors
abstract
Applications executing on multicore processors can now easily offload computations to many core processors, such as Intel Xeon Phi coprocessors. However, it requires high levels of expertise and effort to tune such offloaded applications to realize high-performance execution. Previous efforts have focused on optimizing the execution of offloaded computations on many core processors. However, we observe that the data transfer overhead between multicore and many core processors, and the limited device memories of many core processors often constrain the performance gains that are possible by offloading computations. In this paper, we present three source-to-source compiler optimizations that can significantly improve the performance of applications that offload computations to many core processors. The first optimization automatically transforms offloaded codes to enable data streaming, which overlaps data transfer between multicore and many core processors with computations on these processors to hide data transfer overhead. This optimization is also designed to minimize the memory usage on many core processors, while achieving the optimal performance. The second compiler optimization re-orders computations to regularize irregular memory accesses. It enables data streaming and factorization on many core processors, even when the memory access patterns in the original source codes are irregular. Finally, our new shared memory mechanism provides efficient support for transferring large pointer-based data structures between hosts and many core processors. Our evaluation shows that the proposed compiler optimizations benefit 9 out of 12 benchmarks. Compared with simply offloading the original parallel implementations of these benchmarks, we can achieve 1.16x-52.21x speedups.
Linhai Song, Min Feng 0001, Nishkam Ravi, Yi Yang 0018, Srimat T. Chakradhar
MICRO3
2013 Semi-automatic restructuring of offloadable tasks for many-core accelerators
abstract
Work division between the processor and accelerator is a common theme in modern heterogenous computing. Recent efforts (such as LEO and OpenAcc) provide directives that allow the developer to mark code regions in the original application from which offloadable tasks can be generated by the compiler. Auto-tuners and runtime schedulers work with the options (i.e., offloadable tasks) generated at compile time, which is limited by the directives specified by the developer. There is no provision for offload restructuring.
Nishkam Ravi, Yi Yang 0018, Srimat T. Chakradhar
SC1
2012 Panacea: towards holistic optimization of MapReduce applications
abstract
MapReduce has emerged as one of the most popular programming models for data parallel enterprise applications. Despite advances in runtime, the opportunities for optimizing MapReduce applications remain largely unexplored. In this paper, we present a framework for performing holistic compiler optimizations on legacy MapReduce applications. We have identified and implemented two optimizations and evaluated them with a set of Hadoop applications on a cluster of Xeon servers. Our experiments show that performance gains of more than 3X can be achieved without user involvement.
Jun Liu 0008, Nishkam Ravi, Srimat T. Chakradhar, Mahmut T. Kandemir
CGO2
2012 Apricot: an optimizing compiler and productivity tool for x86-compatible many-core coprocessors
abstract
Intel MIC (Many Integrated Core) is the first x86-based coprocessor architecture aimed at accelerating multi-core HPC applications. In the most common usage model, parallel code sections are offloaded to the MIC coprocessor using LEO (Language Extensions for Offload). The developer is responsible for identifying and specifying offloadable code regions, managing data transfers between the CPU and MIC and optimizing the application for performance, which requires some amount of effort and experimentation. In this paper, we present Apricot, an optimizing compiler and productivity tool for x86-compatible many-core coprocessors (such as Intel MIC) that minimizes developer effort by (i) automatically inserting LEO clauses for parallelizable code regions, (ii) selectively offloading some of the code regions to the coprocessor at runtime based on a cost model that we have developed, (iii) applying a set ofoptimizations for minimizing the data communication overhead and improving overall performance. Apricot is intended to assist programmers in porting existing multi-core applications and writing new ones to take advantage of the many-core coprocessor, while maximizing overall performance. Experiments with SpecOMP and NAS Parallel benchmarks show that Apricot can successfully transform OpenMP applications to run on the MIC coprocessor with good performance gains.
Nishkam Ravi, Yi Yang 0018, Srimat T. Chakradhar
ICS1
2008 Context-aware Battery Management for Mobile Phones
abstract
In this paper, we propose a system for context- aware battery management that warns the user when it detects that the phone battery can run out before the next charging opportunity is encountered. At the heart of this system, are algorithms that predict: (1) when the next charging opportunity will be available, (2) how much call-time will be required by the user in the interim, and (3) how long the battery will last if the current set of applications continue to execute. We propose algorithms that process user's location traces and call-logs for making some of these predictions. We also propose a technique to predict battery consumption of applications. We present the design of the system and demonstrate its feasibility by experimentally showing that each of the prediction algorithms can perform with fairly high accuracy.
Nishkam Ravi, James Scott, Liviu Iftode
PerCom1
2006 Non-Inference: An Information Flow Control Model for Location-based Services
abstract
This paper presents a framework for preserving location privacy without affecting location accuracy. In this framework, services migrate a piece of code to a trusted server, which is assumed to have location information of all the interesting subjects. The code executes on the trusted server, reads location information and sends back results. We introduce Non-inference, a novel information-flow control model that guarantees that the code does not leak exact location information. We discuss the design, implementation and evaluation of a static program analysis technique that enforces non-inference for location based services
Nishkam Ravi, Marco Gruteser, Liviu Iftode
MobiQuitous1
2006 Inverted Browser: A Novel Approach towards Display Symbiosis
abstract
In this paper we introduce the inverted browser, a novel approach to enable mobile users to view content from their personal devices on public displays. The inverted browser is a network service to start and control a browser that is then used to view the content. In contrast to a traditional Web browser, which runs on the client device and pulls content from a server, content is pushed to the inverted browser from a personal data source upon user input. This approach allows a wide variety of personal content to be viewed by facilitating symbiotic relationships between mobile devices and intelligent displays in the environment. Our initial inverted browser prototype is based on a Web services wrapper around a traditional Web browser. Our experiments show that the inverted browser approach is superior to other solutions in terms of user convenience, ease of use, energy consumption, and privacy, but interaction latencies need improvement.
Mandayam T. Raghunath, Nishkam Ravi, Marcel-Catalin Rosu, Chandrasekhar Narayanaswami 0001
PerCom2
2006 An Efficient Optimal-Equilibrium Algorithm for Two-player Game Trees
Michael L. Littman, Nishkam Ravi, Arjun Talwar, Martin Zinkevich
UAI2
2005 Activity Recognition from Accelerometer Data
Nishkam Ravi, Nikhil Dandekar, Preetham Mysore, Michael L. Littman
AAAI1
2005 Accessing Ubiquitous Services Using Smart Phones
abstract
The integration of Bluetooth service discovery protocol (SDP), and GPRS Internet connectivity into phones provides a simple yet powerful infrastructure for accessing services in nomadic environments. In this paper, we discuss the design and implementation of SDIPP, a protocol for provisioning services on smart phones. Although several service discovery protocols have been proposed earlier, such as SLP, Jini, UPnP, Salutation, they all have their own infrastructure requirements and target audiences, Bluetooth SDP is an on-the-fly service discovery protocol. However, it is not nearly as powerful as its counterparts. SDIPP works by augmenting Bluetooth SDP with Web access and personalization. Payment of services has been overlooked in the protocols proposed earlier. SDIPP provides a novel protocol for anonymous payment, based on the idea of Millicent scrips. We have implemented a few services to illustrate our protocol. We report on our experiences and experimental results. In particular, we analyze and provide an application level solution to the Bluetooth inquiry clash problem that was discovered in the process.
Nishkam Ravi, Peter Stern, Niket Desai, Liviu Iftode
PerCom1
2004 An Instance-Based State Representation for Network Repair
Michael L. Littman, Nishkam Ravi, Eitan Fenson, Richard E. Howard
AAAI2
2004 Portable Smart Messages for Ubiquitous Java-Enabled Devices
abstract
Recent advances in wireless technology allow Java-enabled devices, such as Smart Phones and PDAs, to create mobile ad hoc networks, over which distributed applications can be executed. Although Java shields the programmers from the heterogeneity of the hardware platforms, a common middleware architecture is needed to support a cooperative execution environment in these networks. In this paper, we present a portable runtime system for smart messages (SMs), a middleware architecture based on execution migration, that we designed and implemented on top of an unmodified Java virtual machine. To facilitate portability, we have designed a lightweight migration mechanism based on Java bytecode instrumentation. This mechanism is suitable for mobile ad hoc networks where limited bandwidth and mobility impose constraints on the amount of data transferred. The experimental results for applications executed over wireless networks of HP iPAQs demonstrate the feasibility of our portable runtime system.
Nishkam Ravi, Cristian Borcea, Porlin Kang, Liviu Iftode
MobiQuitous1