Krishna Narra

dblp:230/3925 · also H. V. Krishna Giri Narra, Hema Venkata Krishna Giri Narra, Krishna Giri Narra · DBLP profile ↗
← Back
5ranked-venue papers
3as first author
1since 2021 · last 2021
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 2 first-authorArtificial intelligence and machine learning · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
Storage systems · 48% Distributed systems · 32% Cloud and datacenter computing · 20%
Artificial intelligence
1 paper
Efficient and distributed learning · 100%

Topics — the 12 heaviest of 12, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Distributed systems
coded computation
0.412019
Slack squeeze coded computing for adaptive straggler mitigation · SC 2019
Distributed systems › distributed data processing
straggler mitigation
0.412019
Slack squeeze coded computing for adaptive straggler mitigation · SC 2019
Cloud and datacenter computing › resource allocation
workload allocation
0.412019
Slack squeeze coded computing for adaptive straggler mitigation · SC 2019
Machine learning › Efficient and distributed learning › distributed training
communication-efficient training
0.312018
GradiVeQ: Vector Quantization for Bandwidth-Efficient Gradient Aggregation in Distributed CNN Training · NeurIPS 2018
Machine learning › Efficient and distributed learning
distributed training
0.312018
GradiVeQ: Vector Quantization for Bandwidth-Efficient Gradient Aggregation in Distributed CNN Training · NeurIPS 2018
Machine learning › Efficient and distributed learning › distributed training
gradient compression
0.312018
GradiVeQ: Vector Quantization for Bandwidth-Efficient Gradient Aggregation in Distributed CNN Training · NeurIPS 2018
Machine learning › Efficient and distributed learning › communication compression
gradient quantization
0.312018
GradiVeQ: Vector Quantization for Bandwidth-Efficient Gradient Aggregation in Distributed CNN Training · NeurIPS 2018
Storage systems › flash and SSD › flash memory management
flash translation layer
0.312017
Summarizer: trading communication with computing near storage · MICRO 2017
Storage systems › computational storage
in-storage computing
0.312017
Summarizer: trading communication with computing near storage · MICRO 2017
Storage systems › computational storage
near-storage computing
0.312017
Summarizer: trading communication with computing near storage · MICRO 2017
Storage systems › flash and SSD
solid-state drive
0.312017
Summarizer: trading communication with computing near storage · MICRO 2017
Cloud and datacenter computing
datacenter storage
0.112017
Summarizer: trading communication with computing near storage · MICRO 2017

Methods — techniques the papers use, named apart from their topics

coding theory · 0.4LSTM-based speed prediction · 0.4vector quantization · 0.3ring all-reduce · 0.3principal component analysis · 0.3embedded core offloading · 0.3
YearPublicationVenuePosition
2021 Origami Inference: Private Inference Using Hardware Enclaves
abstract
This work presents Origami, a framework which provides privacy-preserving inference for large deep neural network (DNN) models through a combination of enclave execution, cryptographic blinding, interspersed with accelerator-based computation. Origami partitions the ML model into multiple partitions. The first partition receives the encrypted user input within an SGX enclave. The enclave decrypts the input and then applies cryptographic blinding to the input data and the model parameters. The layer computation is offloaded to a GPU/CPU and the computed output is returned to the enclave, which decodes the computation on noisy data using the unblinding factors privately stored within SGX. This process may be repeated for each DNN layer, as has been done in prior work Slalom. However, the overhead of blinding and unblinding the data is a limiting factor to scalability. Origami relies on the empirical observation that the feature maps after the first several layers can not be used, even by a powerful conditional GAN adversary to reconstruct input. Hence, Origami dynamically switches to executing the rest of the DNN layers directly on an accelerator. We empirically demonstrate that using Origami, a conditional GAN adversary, even with an unlimited inference budget, cannot reconstruct the input. Compared to running the entire VGG-19 model within SGX, Origami inference improves the performance of private inference from 11x while using Slalom to 15. 1x.
Krishna Narra, Zhifeng Lin, Yongqin Wang, Keshav Balasubramanian, Murali Annavaram
CLOUD1
2020 Collage Inference: Using Coded Redundancy for Lowering Latency Variation in Distributed Image Classification Systems
abstract
MLaaS (ML-as-a-Service) offerings by cloud computing platforms are becoming increasingly popular. Hosting pre-trained machine learning models in the cloud enables elastic scalability as the demand grows. But providing low latency and reducing the latency variance is a key requirement. Variance is harder to control in a cloud deployment due to uncertain-ties in resource allocations across many virtual instances. We propose the collage inference technique, which uses a novel convolutional neural network model, collage-cnn, to provide low-cost redundancy. A collage-cnn model takes a collage image formed by combining multiple images and performs multi-image classification in one shot, albeit at slightly lower accuracy. We augment a collection of traditional single image classifier models with a single collage-cnn classifier, which acts as their low-cost redundant backup. Collage-cnn provides backup classification results if any single image classification requests experience a slowdown. Deploying the collage-cnn models in the cloud, we demonstrate that the 99th percentile tail latency of inference can be reduced by 1.2x to 2x compared to replication-based approaches while providing high accuracy. Variation in inference latency can be reduced by 1.8x to 15x.
Krishna Narra, Zhifeng Lin, Ganesh Ananthanarayanan, Amir Salman Avestimehr, Murali Annavaram
ICDCS1
2019 Slack squeeze coded computing for adaptive straggler mitigation
abstract
While performing distributed computations in today's cloud-based platforms, execution speed variations among compute nodes can significantly reduce the performance and create bottlenecks like stragglers. Coded computation techniques leverage coding theory to inject computational redundancy and mitigate stragglers in distributed computations. In this paper, we propose a dynamic workload distribution strategy for coded computation called Slack Squeeze Coded Computation (S2C2). S2C2 squeezes the compute slack (i.e., overhead) that is built into the coded computing frameworks by efficiently assigning work for all fast and slow nodes according to their speeds and without needing to re-distribute data. We implement an LSTM-based speed prediction algorithm to predict speeds of compute nodes. We evaluate S2C2 on linear algebraic algorithms, gradient descent, graph ranking, and graph filtering algorithms. We demonstrate 19% to 39% reduction in total computation latency using S2C2 compared to job replication and coded computation. We further show how S2C2 can be applied beyond matrix-vector multiplication.
Krishna Narra, Zhifeng Lin, Mehrdad Kiamari, Amir Salman Avestimehr, Murali Annavaram
SC1
2018 GradiVeQ: Vector Quantization for Bandwidth-Efficient Gradient Aggregation in Distributed CNN Training
abstract
Data parallelism can boost the training speed of convolutional neural networks (CNN), but could suffer from significant communication costs caused by gradient aggregation. To alleviate this problem, several scalar quantization techniques have been developed to compress the gradients. But these techniques could perform poorly when used together with decentralized aggregation protocols like ring all-reduce (RAR), mainly due to their inability to directly aggregate compressed gradients. In this paper, we empirically demonstrate the strong linear correlations between CNN gradients, and propose a gradient vector quantization technique, named GradiVeQ, to exploit these correlations through principal component analysis (PCA) for substantial gradient dimension reduction. GradiveQ enables direct aggregation of compressed gradients, hence allows us to build a distributed learning system that parallelizes GradiveQ gradient compression and RAR communications. Extensive experiments on popular CNNs demonstrate that applying GradiveQ slashes the wall-clock gradient aggregation time of the original RAR by more than 5x without noticeable accuracy loss, and reduce the end-to-end training time by almost 50%. The results also show that \GradiveQ is compatible with scalar quantization techniques such as QSGD (Quantized SGD), and achieves a much higher speed-up gain under the same compression ratio.
Mingchao Yu, Zhifeng Lin, Krishna Narra, Youjie Li, Nam Sung Kim, Alexander G. Schwing, Murali Annavaram, Amir Salman Avestimehr
NeurIPS3
2017 Summarizer: trading communication with computing near storage
abstract
Modern data center solid state drives (SSDs) integrate multiple general-purpose embedded cores to manage flash translation layer, garbage collection, wear-leveling, and etc., to improve the performance and the reliability of SSDs. As the performance of these cores steadily improves there are opportunities to repurpose these cores to perform application driven computations on stored data, with the aim of reducing the communication between the host processor and the SSD. Reducing host-SSD bandwidth demand cuts down the I/O time which is a bottleneck for many applications operating on large data sets. However, the embedded core performance is still significantly lower than the host processor, as generally wimpy embedded cores are used within SSD for cost effective reasons. So there is a trade-off between the computation overhead associated with near SSD processing and the reduction in communication overhead to the host system.
Gunjae Koo, Kiran Kumar Matam, Te I, Krishna Narra, Jing Li 0021, Hung-Wei Tseng 0001, Steven Swanson, Murali Annavaram
MICRO4