EDBT 2026 Demo / reviewers in the wild / expert
Vikas Natesh
dblp:246/9390
· DBLP profile ↗
3ranked-venue papers
2as first author
3since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 3 · 2 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MGS: Markov Greedy Sums for Low-Power DNN Accumulation
Vikas Natesh, H. T. Kung 0001 |
IPDPS | 1 |
| 2025 | Alternating Greedy Schedules: Enabling Low-Bitwidth Accumulation of Dot Products in Neural Network ComputationsabstractWe present Alternating Greedy Scheduling (AGS), an algorithm for avoiding overflow, specifically transient overflow, during low-bitwidth accumulation of dot products in neural network computations. In conventional quantized (e.g., 8-bit) dot products, partial results are accumulated into wide (e.g., 32-bit) accumulators to avoid overflows when accumulating intermediate partial sums. However, such wide accumulators increase memory bandwidth usage and reduce energy efficiency. We show that iterative N:M pruning in floating point followed by quantization to 8 (or fewer) bits, and accumulation of partial products in an optimal order (via AGS) allows for accurate, compressed models with a large number of partial products that do not require wide accumulators. We design, analyze, and implement the AGS algorithm to eliminate accumulation overflows at inference time for several neural networks. Our method offers a 2.7x reduction in accumulator bitwidth while achieving model accuracy on par with floating-point baselines for multiple image classification tasks. Vikas Natesh, H. T. Kung 0001 |
ISCAS | 1 |
| 2021 | CAKE: matrix multiplication using constant-bandwidth blocksabstractWe offer a novel approach to matrix-matrix multiplication computation on computing platforms with memory hierarchies. Constant-bandwidth (CB) blocks improve computation throughput for architectures limited by external memory bandwidth. Configuring the shape and size of CB blocks operating from within any memory hierarchy level (e.g., internal SRAM), we achieve high throughput while holding external bandwidth (e.g., with DRAM) constant. We explain how, surprisingly, CB blocks can maintain constant external bandwidth as computation throughput increases. Analogous to partitioning a cake into pieces, we dub our CB-partitioned system CAKE. H. T. Kung 0001, Vikas Natesh, Andrew Sabot |
SC | 2 |