EDBT 2026 Demo / reviewers in the wild / expert
Krishna Narra
dblp:230/3925 · also H. V. Krishna Giri Narra, Hema Venkata Krishna Giri Narra, Krishna Giri Narra
· DBLP profile ↗
5ranked-venue papers
3as first author
1since 2021 · last 2021
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 3 · 2 first-authorArtificial intelligence and machine learning · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
2 papers |
Storage systems · 48% Distributed systems · 32% Cloud and datacenter computing · 20% | |
| Artificial intelligence
1 paper |
Efficient and distributed learning · 100% |
Topics — the 12 heaviest of 12, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Distributed systems
coded computation |
0.4 | 1 | 2019 | Slack squeeze coded computing for adaptive straggler mitigation · SC 2019 |
Distributed systems › distributed data processing
straggler mitigation |
0.4 | 1 | 2019 | Slack squeeze coded computing for adaptive straggler mitigation · SC 2019 |
Cloud and datacenter computing › resource allocation
workload allocation |
0.4 | 1 | 2019 | Slack squeeze coded computing for adaptive straggler mitigation · SC 2019 |
Machine learning › Efficient and distributed learning › distributed training
communication-efficient training |
0.3 | 1 | 2018 | GradiVeQ: Vector Quantization for Bandwidth-Efficient Gradient Aggregation in Distributed CNN Training · NeurIPS 2018 |
Machine learning › Efficient and distributed learning
distributed training |
0.3 | 1 | 2018 | GradiVeQ: Vector Quantization for Bandwidth-Efficient Gradient Aggregation in Distributed CNN Training · NeurIPS 2018 |
Machine learning › Efficient and distributed learning › distributed training
gradient compression |
0.3 | 1 | 2018 | GradiVeQ: Vector Quantization for Bandwidth-Efficient Gradient Aggregation in Distributed CNN Training · NeurIPS 2018 |
Machine learning › Efficient and distributed learning › communication compression
gradient quantization |
0.3 | 1 | 2018 | GradiVeQ: Vector Quantization for Bandwidth-Efficient Gradient Aggregation in Distributed CNN Training · NeurIPS 2018 |
Storage systems › flash and SSD › flash memory management
flash translation layer |
0.3 | 1 | 2017 | Summarizer: trading communication with computing near storage · MICRO 2017 |
Storage systems › computational storage
in-storage computing |
0.3 | 1 | 2017 | Summarizer: trading communication with computing near storage · MICRO 2017 |
Storage systems › computational storage
near-storage computing |
0.3 | 1 | 2017 | Summarizer: trading communication with computing near storage · MICRO 2017 |
Storage systems › flash and SSD
solid-state drive |
0.3 | 1 | 2017 | Summarizer: trading communication with computing near storage · MICRO 2017 |
Cloud and datacenter computing
datacenter storage |
0.1 | 1 | 2017 | Summarizer: trading communication with computing near storage · MICRO 2017 |
Methods — techniques the papers use, named apart from their topics
coding theory · 0.4LSTM-based speed prediction · 0.4vector quantization · 0.3ring all-reduce · 0.3principal component analysis · 0.3embedded core offloading · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2021 | Origami Inference: Private Inference Using Hardware EnclavesabstractThis work presents Origami, a framework which provides privacy-preserving inference for large deep neural network (DNN) models through a combination of enclave execution, cryptographic blinding, interspersed with accelerator-based computation. Origami partitions the ML model into multiple partitions. The first partition receives the encrypted user input within an SGX enclave. The enclave decrypts the input and then applies cryptographic blinding to the input data and the model parameters. The layer computation is offloaded to a GPU/CPU and the computed output is returned to the enclave, which decodes the computation on noisy data using the unblinding factors privately stored within SGX. This process may be repeated for each DNN layer, as has been done in prior work Slalom. However, the overhead of blinding and unblinding the data is a limiting factor to scalability. Origami relies on the empirical observation that the feature maps after the first several layers can not be used, even by a powerful conditional GAN adversary to reconstruct input. Hence, Origami dynamically switches to executing the rest of the DNN layers directly on an accelerator. We empirically demonstrate that using Origami, a conditional GAN adversary, even with an unlimited inference budget, cannot reconstruct the input. Compared to running the entire VGG-19 model within SGX, Origami inference improves the performance of private inference from 11x while using Slalom to 15. 1x. Krishna Narra, Zhifeng Lin, Yongqin Wang, Keshav Balasubramanian, Murali Annavaram |
CLOUD | 1 |
| 2020 | Collage Inference: Using Coded Redundancy for Lowering Latency Variation in Distributed Image Classification SystemsabstractMLaaS (ML-as-a-Service) offerings by cloud computing platforms are becoming increasingly popular. Hosting pre-trained machine learning models in the cloud enables elastic scalability as the demand grows. But providing low latency and reducing the latency variance is a key requirement. Variance is harder to control in a cloud deployment due to uncertain-ties in resource allocations across many virtual instances. We propose the collage inference technique, which uses a novel convolutional neural network model, collage-cnn, to provide low-cost redundancy. A collage-cnn model takes a collage image formed by combining multiple images and performs multi-image classification in one shot, albeit at slightly lower accuracy. We augment a collection of traditional single image classifier models with a single collage-cnn classifier, which acts as their low-cost redundant backup. Collage-cnn provides backup classification results if any single image classification requests experience a slowdown. Deploying the collage-cnn models in the cloud, we demonstrate that the 99th percentile tail latency of inference can be reduced by 1.2x to 2x compared to replication-based approaches while providing high accuracy. Variation in inference latency can be reduced by 1.8x to 15x. Krishna Narra, Zhifeng Lin, Ganesh Ananthanarayanan, Amir Salman Avestimehr, Murali Annavaram |
ICDCS | 1 |
| 2019 | Slack squeeze coded computing for adaptive straggler mitigationabstractWhile performing distributed computations in today's cloud-based platforms, execution speed variations among compute nodes can significantly reduce the performance and create bottlenecks like stragglers. Coded computation techniques leverage coding theory to inject computational redundancy and mitigate stragglers in distributed computations. In this paper, we propose a dynamic workload distribution strategy for coded computation called Slack Squeeze Coded Computation (S2C2). S2C2 squeezes the compute slack (i.e., overhead) that is built into the coded computing frameworks by efficiently assigning work for all fast and slow nodes according to their speeds and without needing to re-distribute data. We implement an LSTM-based speed prediction algorithm to predict speeds of compute nodes. We evaluate S2C2 on linear algebraic algorithms, gradient descent, graph ranking, and graph filtering algorithms. We demonstrate 19% to 39% reduction in total computation latency using S2C2 compared to job replication and coded computation. We further show how S2C2 can be applied beyond matrix-vector multiplication. Krishna Narra, Zhifeng Lin, Mehrdad Kiamari, Amir Salman Avestimehr, Murali Annavaram |
SC | 1 |
| 2018 | GradiVeQ: Vector Quantization for Bandwidth-Efficient Gradient Aggregation in Distributed CNN TrainingabstractData parallelism can boost the training speed of convolutional neural networks (CNN), but could suffer from significant communication costs caused by gradient aggregation. To alleviate this problem, several scalar quantization techniques have been developed to compress the gradients. But these techniques could perform poorly when used together with decentralized aggregation protocols like ring all-reduce (RAR), mainly due to their inability to directly aggregate compressed gradients. In this paper, we empirically demonstrate the strong linear correlations between CNN gradients, and propose a gradient vector quantization technique, named GradiVeQ, to exploit these correlations through principal component analysis (PCA) for substantial gradient dimension reduction. GradiveQ enables direct aggregation of compressed gradients, hence allows us to build a distributed learning system that parallelizes GradiveQ gradient compression and RAR communications. Extensive experiments on popular CNNs demonstrate that applying GradiveQ slashes the wall-clock gradient aggregation time of the original RAR by more than 5x without noticeable accuracy loss, and reduce the end-to-end training time by almost 50%. The results also show that \GradiveQ is compatible with scalar quantization techniques such as QSGD (Quantized SGD), and achieves a much higher speed-up gain under the same compression ratio. Mingchao Yu, Zhifeng Lin, Krishna Narra, Youjie Li, Nam Sung Kim, Alexander G. Schwing, Murali Annavaram, Amir Salman Avestimehr |
NeurIPS | 3 |
| 2017 | Summarizer: trading communication with computing near storageabstractModern data center solid state drives (SSDs) integrate multiple general-purpose embedded cores to manage flash translation layer, garbage collection, wear-leveling, and etc., to improve the performance and the reliability of SSDs. As the performance of these cores steadily improves there are opportunities to repurpose these cores to perform application driven computations on stored data, with the aim of reducing the communication between the host processor and the SSD. Reducing host-SSD bandwidth demand cuts down the I/O time which is a bottleneck for many applications operating on large data sets. However, the embedded core performance is still significantly lower than the host processor, as generally wimpy embedded cores are used within SSD for cost effective reasons. So there is a trade-off between the computation overhead associated with near SSD processing and the reduction in communication overhead to the host system. Gunjae Koo, Kiran Kumar Matam, Te I, Krishna Narra, Jing Li 0021, Hung-Wei Tseng 0001, Steven Swanson, Murali Annavaram |
MICRO | 4 |