Kiran Kumar Matam

dblp:28/9586 · DBLP profile ↗
← Back
12ranked-venue papers
7as first author
4since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 10 · 6 first-author · 2 since 2021Computer networks · 2 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
5 papers
Storage systems · 36% Hardware accelerators and domain-specific architectures · 31% Distributed systems · 27%
Databases, data mining, and information retrieval
3 papers
Recommender systems · 73% Machine learning and data management · 18% Graph data management · 9%
Artificial intelligence
2 papers
Efficient and distributed learning · 100%

Topics — the 15 heaviest of 18, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Recommender systems › web personalization
real-time personalization
0.812024
QuickUpdate: a Real-Time Personalization System for Large-Scale Recommendation Models · NSDI 2024
Machine learning › Efficient and distributed learning
distributed training
0.722022
Software-hardware co-design for fast and scalable training of deep learning recommendation models · ISCA 2022
Check-N-Run: a Checkpointing System for Training Deep Learning Recommendation Models · NSDI 2022
Distributed systems › fault tolerance
checkpointing
0.612022
Check-N-Run: a Checkpointing System for Training Deep Learning Recommendation Models · NSDI 2022
Distributed systems
fault tolerance
0.612022
Check-N-Run: a Checkpointing System for Training Deep Learning Recommendation Models · NSDI 2022
Hardware accelerators and domain-specific architectures
machine learning accelerator
0.612022
Software-hardware co-design for fast and scalable training of deep learning recommendation models · ISCA 2022
Hardware accelerators and domain-specific architectures › machine learning accelerator
recommendation model training
0.612022
Software-hardware co-design for fast and scalable training of deep learning recommendation models · ISCA 2022
Storage systems
flash and SSD
0.412019
GraphSSD: graph semantics aware SSD · ISCA 2019
Storage systems › flash and SSD › flash memory management
flash translation layer
0.312017
Summarizer: trading communication with computing near storage · MICRO 2017
Storage systems › computational storage
in-storage computing
0.312017
Summarizer: trading communication with computing near storage · MICRO 2017
Storage systems › computational storage
near-storage computing
0.312017
Summarizer: trading communication with computing near storage · MICRO 2017
Storage systems › flash and SSD
solid-state drive
0.312017
Summarizer: trading communication with computing near storage · MICRO 2017
Machine learning and data management
inference serving
0.212024
QuickUpdate: a Real-Time Personalization System for Large-Scale Recommendation Models · NSDI 2024
Graph data management › graph processing
out-of-core graph processing
0.112019
GraphSSD: graph semantics aware SSD · ISCA 2019
Cloud and datacenter computing
datacenter storage
0.112017
Summarizer: trading communication with computing near storage · MICRO 2017
Network measurement and analytics › traffic classification
internet traffic classification
0.012013
High throughput and programmable online trafficclassifier on FPGA · FPGA 2013

Methods — techniques the papers use, named apart from their topics

software-managed caching · 1.7kernel fusion · 1.7embedding compression · 1.7NVMe interface extension · 0.8entropy-MDL discretization · 0.3c4.5 decision tree · 0.3embedded core offloading · 0.3
YearPublicationVenuePosition
2024 QuickUpdate: a Real-Time Personalization System for Large-Scale Recommendation Models
Kiran Kumar Matam, Hani Ramezani, Zeliang Chen, Maomao Ding, Ellie Wen, Assaf Eisenman
NSDI1
2022 Software-hardware co-design for fast and scalable training of deep learning recommendation models
abstract
Deep learning recommendation models (DLRMs) have been used across many business-critical services at Meta and are the single largest AI application in terms of infrastructure demand in its data-centers. In this paper, we present Neo, a software-hardware co-designed system for high-performance distributed training of large-scale DLRMs. Neo employs a novel 4D parallelism strategy that combines table-wise, row-wise, column-wise, and data parallelism for training massive embedding operators in DLRMs. In addition, Neo enables extremely high-performance and memory-efficient embedding computations using a variety of critical systems optimizations, including hybrid kernel fusion, software-managed caching, and quality-preserving compression. Finally, Neo is paired with ZionEX, a new hardware platform co-designed with Neo's 4D parallelism for optimizing communications for large-scale DLRM training. Our evaluation on 128 GPUs using 16 ZionEX nodes shows that Neo outperforms existing systems by up to 40× for training 12-trillion-parameter DLRM models deployed in production.
Dheevatsa Mudigere, Yuchen Hao, Andrew Tulloch, Srinivas Sridharan 0002, Muhammet Mustafa Ozdal, Jade Nie, Jongsoo Park, Jie Amy Yang, Leon Gao, Dmytro Ivchenko, Aarti Basant, Yuxi Hu 0001, Jiyan Yang, Ehsan K. Ardestani, Xiaodong Wang 0020, Rakesh Komuravelli, Ching-Hsiang Chu, Serhat Yilmaz, Jiyuan Qian, Zhuobo Feng, Yinbin Ma, Junjie Yang 0005, Ellie Wen, Chonglin Sun, Whitney Zhao, Dimitry Melts, Krishna Dhulipala, K. R. Kishore, Tyler Graf, Assaf Eisenman, Kiran Kumar Matam, Adi Gangidi, Guoqiang Jerry Chen, Manoj Krishnan, Avinash Nayak, Krishnakumar Nair, Bharath Muthiah, Mahmoud khorashadi, Pallab Bhattacharya, Petr Lapukhov, Maxim Naumov, Ajit Mathews, Lin Qiao, Mikhail Smelyanskiy, Bill Jia, Vijay Rao
ISCA38
2022 Check-N-Run: a Checkpointing System for Training Deep Learning Recommendation Models
Assaf Eisenman, Kiran Kumar Matam, Steven Ingram, Dheevatsa Mudigere, Raghuraman Krishnamoorthi, Krishnakumar Nair, Mikhail Smelyanskiy, Murali Annavaram
NSDI2
2021 MultiLogVC: Efficient Out-of-Core Graph Processing Framework for Flash Storage
abstract
Graph analytics are at the heart of a broad range of applications such as drug discovery, page ranking, transportation systems, and recommendation models. When graph size exceeds the available memory size in a computing node, out-of-core graph processing is needed. For the widely used out-of-core graph processing systems, the graphs are stored and accessed from a long latency SSD storage, which becomes a significant performance bottleneck. To tackle this long latency this work exploits the key insight that that nearly all graph algorithms have a dynamically varying number of active vertices that must be processed in each iteration. However, existing graph processing frameworks, such as GraphChi, load the entire graph in each iteration even if a small fraction of the graph is active. This limitation is due to the structure of the graph storage used by these systems. In this work, we propose to use a compressed sparse row (CSR) based graph storage that is more amenable for selectively loading only a few active vertices in each iteration. However, CSR based graph processing suffers from random update propagation to many target vertices. To solve this challenge, we propose to use a multi-log update mechanism that logs updates separately, rather than directly update the active edges and vertices in a graph. The multi-log system maintains a separate log per each vertex interval (a group of vertices). This separation enables efficient processing of all updates bound to each vertex interval by just loading the corresponding log. Further, by logging all the updates associated with a vertex interval in one contiguous log this approach reduces read amplification since all the pages in the log will be processed in the next iteration without wasted page reads. Over the current state of the art out-of-core graph processing framework, our evaluation results show that the MultiLogVC framework improves performance by up to 17.84×, 1.19×, 1.65×, 1.38×, 3.15×, and 6.00× for the widely used breadth-first search, pagerank, community detection, graph coloring, maximal independent set, and random-walk applications, respectively.
Kiran Kumar Matam, Hanieh Hashemi, Murali Annavaram
IPDPS1
2019 GraphSSD: graph semantics aware SSD
abstract
Graph analytics play a key role in a number of applications such as social networks, drug discovery, and recommendation systems. Given the large size of graphs that may exceed the capacity of the main memory, application performance is bounded by storage access time. Out-of-core graph processing frameworks try to tackle this storage access bottleneck through techniques such as graph sharding, and sub-graph partitioning. Even with these techniques, the need to access data across different graph shards or sub-graphs causes storage systems to become a significant performance hurdle. In this paper, we propose a graph semantic aware solid state drive (SSD) framework, called GraphSSD, which is a full system solution for storing, accessing, and performing graph analytics on SSDs. Rather than treating storage as a collection of blocks, GraphSSD considers graph structure while deciding on graph layout, access, and update mechanisms. GraphSSD replaces the conventional logical to physical page mapping mechanism in an SSD with a novel vertex-to-page mapping scheme and exploits the detailed knowledge of the flash properties to minimize page accesses. GraphSSD also supports efficient graph updates (vertex and edge modifications) by minimizing unnecessary page movement overheads. GraphSSD provides a simple programming interface that enables application developers to access graphs as native data in their applications, thereby simplifying the code development. It also augments the NVMe (non-volatile memory express) interface with a minimal set of changes to map the graph access APIs to appropriate storage access mechanisms.
Kiran Kumar Matam, Gunjae Koo, Haipeng Zha 0001, Hung-Wei Tseng 0001, Murali Annavaram
ISCA1
2019 Efficient automatic parallelization of a single GPU program for a multiple GPU system
Kiran Kumar Matam, Mohammad Abdel-Majeed, Murali Annavaram
Integr.1
2017 Summarizer: trading communication with computing near storage
abstract
Modern data center solid state drives (SSDs) integrate multiple general-purpose embedded cores to manage flash translation layer, garbage collection, wear-leveling, and etc., to improve the performance and the reliability of SSDs. As the performance of these cores steadily improves there are opportunities to repurpose these cores to perform application driven computations on stored data, with the aim of reducing the communication between the host processor and the SSD. Reducing host-SSD bandwidth demand cuts down the I/O time which is a bottleneck for many applications operating on large data sets. However, the embedded core performance is still significantly lower than the host processor, as generally wimpy embedded cores are used within SSD for cost effective reasons. So there is a trade-off between the computation overhead associated with near SSD processing and the reduction in communication overhead to the host system.
Gunjae Koo, Kiran Kumar Matam, Te I, Krishna Narra, Jing Li 0021, Hung-Wei Tseng 0001, Steven Swanson, Murali Annavaram
MICRO2
2013 High throughput and programmable online trafficclassifier on FPGA
abstract
Machine learning (ML) algorithms have been shown to be effective in classifying the dynamic internet traffic today. Using additional features and sophisticated ML techniques can improve accuracy and can classify a broad range of application classes. Realizing such classifiers to meet high data rates is challenging. In this paper, we propose two architectures to realize complete online traffic classifier using flow-level features. First, we develop a traffic classifier based on C4.5 decision tree algorithm and Entropy-MDL discretization algorithm. It achieves an accuracy of 97.92% when classifying a traffic trace consisting of eight application classes. Next, we accelerate our classifier using two architectures on FPGA. One architecture stores the classifier in on-chip distributed RAM. It is designed to sustain a high throughput. The other architecture stores the classifier in block RAM. It is designed to operate with small hardware footprint and thus built at low hardware cost. Experimental results show that our high throughput architecture can sustain a throughput of $550$ Gbps assuming 40 Byte packet size. Our low cost architecture demonstrates a 22% better resource efficiency than the high throughput design. It can be easily replicated to achieve $449$ Gbps while supporting 160 input traffic streams concurrently. Both architectures are parameterizable and programmable to support any binary-tree-based traffic classifier. We develop a tool which allows users to easily map a binary-tree-based classifier to hardware. The tool takes a classifier as input and automatically generates the Verilog code for the corresponding hardware architecture.
Da Tong, Kiran Kumar Matam, Viktor Prasanna 0001
FPGA3
2013 Energy efficient architecture for matrix multiplication on FPGAs
abstract
Energy efficiency has emerged as one of the key performance metrics. In this work, we first implement a baseline architecture for matrix multiplication, parameterized with the number of processing elements and the types of storage memory. We map this architecture onto a state-of-the-art Field Programmable Gate Array (FPGA). A design space is generated to demonstrate the effect of these parameters on the energy efficiency (defined as number of operations per Joule). We determine that on-chip memory constitutes the largest amount of power consumption among all the components. To improve energy performance, we propose a memory activation schedule. Using this scheme, the proposed optimized design achieves 2.2x and 1.33x improvement with respect to Energy×Area×Time (EAT) and energy efficiency, respectively, compared with the state-of-the-art matrix multiplication core.
Kiran Kumar Matam, Hoang Le, Viktor Prasanna 0001
FPL1
2012 Sparse matrix-matrix multiplication on modern architectures
abstract
Sparse matrix-sparse/dense matrix multiplications, spgemm and csrmm, respectively, among other applications find usage in various matrix formulations of graph problems. Considering the difficulties in executing graph problems and the duality between graphs and matrices, computations such as spgemm and csrmm have recently caught the attention of HPC community. These computations pose challenges such as load balancing, irregular nature of the computation, and difficulty in predicting the output size. It is even more challenging when combined with the GPU architectural constraints such as memory accesses, limited shared memory, strict SIMD and thread execution. To address these challenges on a GPU, we evaluate three possible variations of matrix multiplication (Row-Column, Column-Row, Row-Row) and perform suitable optimizations targeted at sparse matrices. Our experiments indicate that the Row-Row formulation, which mostly outperforms the other formulations, is 3.5x faster on average compared to an optimized multi-core implementation in the Intel MKL library. We extend the Row-Row formulation to a CPU+GPU hybrid algorithm that simultaneously utilizes the CPU also. In this direction, we present heuristics to find the right amount of work division between the CPU and the GPU. Our hybrid row-row formulation of the spgemm operation performs 5.5x faster on average when compared to the optimized multi-core implementation in the Intel MKL library. Our experience indicates that it is difficult to identify right amount of work division between the CPU and the GPU. We therefore investigate a subclass of sparse matrices, band matrices, and present an analytical method to identify a good work division when multiplying two band matrices. Our GPU csrmm operation performs 2.5x faster on average when compared to a corresponding implementation in the cusparse library, which outperforms the Intel MKL library implementation.
Kiran Kumar Matam, Siva Rama Krishna Bharadwaj Indarapu, Kishore Kothapalli
HiPC1
2011 Accelerating Sparse Matrix Vector Multiplication in Iterative Methods Using GPU
abstract
Multiplying a sparse matrix with a vector (spmv for short) is a fundamental operation in many linear algebra kernels. Having an efficient spmv kernel on modern architectures such as the GPUs is therefore of principal interest. The computational challenges that spmv poses are significantlydifferent compared to that of the dense linear algebra kernels. Recent work in this direction has focused on designing data structures to represent sparse matrices so as to improve theefficiency of spmv kernels. However, as the nature of sparseness differs across sparse matrices, there is no clear answer as to which data structure to use given a sparse matrix. In this work, we address this problem by devising techniques to understand the nature of the sparse matrix and then choose appropriate data structures accordingly. By using our technique, we are able to improve the performance of the spmv kernel on an Nvidia Tesla GPU (C1060) by a factor of up to80% in some instances, and about 25% on average compared to the best results of Bell and Garland [3] on the standard dataset (cf. Williams et al. SC'07) used in recent literature. We also use our spmv in the conjugate gradient method and show an average 20% improvement compared to using HYB spmv of [3], on the dataset obtained from the The University of Florida Sparse Matrix Collection [9].
Kiran Kumar Matam, Kishore Kothapalli
ICPP1
2010 Efficient Discrete Range Searching primitives on the GPU with applications
abstract
Graphics processing units provide a large computational power at a very low price which position them as an ubiquitous accelerator. Efficient primitives that can expand the r ange of operations performed on the GPU are thus important. Discrete Range Searching(DRS) is one such primitive with direct applications to string processing, document and text retrieval systems, and least common ancestor queries. In this work, we present a GPU specific implementation of DRS with an optimal space-time trade off. Toward this end, we also present GPU amenable succinct representations and discuss limitations on the GPU. Our method uses 7.5 bits of additional space per element. The speedup achieved by our method is in the range of 20-25 for preprocessing, and 25-35 for batch querying over a sequential implementation. Compared to an 8-threaded implementation, our methods obtain a speedup of 6-8. We study applications of the DRS on the GPU. Also, we suggest that most graph algorithms which focus on using least common ancestor, can easily be enabled on the GPU based on range minima primitive. Beyond this, we show applications of DRS in string querying and tree queries, and suggest how DRS can be helpful in implementing tree based graph algorithms on the GPU.
Jyothish Soman, Kiran Kumar Matam, Kishore Kothapalli, P. J. Narayanan
HiPC2