Meena Arunachalam

dblp:172/2797 · DBLP profile ↗
← Back
5ranked-venue papers
0as first author
4since 2021 · last 2025
0000-0002-3155-6269ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 3 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021
YearPublicationVenuePosition
2025 TPNM: A CXL Based General Purpose Tiered Process Near Memory Framework
abstract
Near-Memory Processing (NMP) has gained significant attention for its potential to accelerate various workloads. However, NMP performance suffers from challenges related to data locality and scalability, particularly in disaggregated datacenter environments. To address these issues, this paper presents TPNM (Tiered Processing Near Memory), a novel framework for near-data processing in disaggregated memory settings. TPNM leverages Compute Express Link (CXL) technology to enable processing both within memory devices and at the fabric switch level, creating a tiered approach to data processing. Evaluations across diverse workloads demonstrate significant performance improvements over baseline and existing near-data processing approaches. In particular, TPNM reduces latency by up to 4.76x compared to CPU-only baselines and by as much as 62 % compared to existing NMP-based solutions.
Pingyi Huo, Anusha Devulapally, Hasan Al Maruf, Meena Arunachalam, Mahmut T. Kandemir, Narayanan Vijaykrishnan
ISPASS4
2025 Multi-Dimensional ML-Pipeline Optimization in Cost-Effective Disaggregated Datacenter
Pingyi Huo, Anusha Devulapally, Hasan Al Maruf, Nandhini Chandramoorthy, Meena Arunachalam, Gulsum Gudukbay Akbulut, Mahmut T. Kandemir, Narayanan Vijaykrishnan
MICRO5
2024 PIFS-Rec: Process-In-Fabric-Switch for Large-Scale Recommendation System Inferences
abstract
Deep Learning Recommendation Models (DLRMs) have become increasingly popular and prevalent in today's datacenters, consuming most of the AI inference cycles. The performance of DLRMs is heavily influenced by available band-width due to their large vector sizes in embedding tables and concurrent accesses. To achieve substantial improvements over existing solutions, novel approaches towards DLRM optimization are needed, especially, in the context of emerging interconnect technologies like CXL. This study delves into exploring CXL-enabled systems, implementing a process-in-fabric-switch (PIFS) solution to accelerate DLRMs while optimizing their memory and bandwidth scalability. We present an in-depth characterization of industry-scale DLRM workloads running on CXL-ready systems, identifying the predominant bottlenecks in existing CXL systems. We, therefore, propose PIFS-Rec, a PIFS-based scheme that implements near-data processing through downstream ports of the fabric switch. PIFS-Rec achieves a latency that is 3.89 x lower than Pond, an industry-standard CXL-based system, and also outperforms BEACON, a state-of-the-art scheme, by 2.03x.
Pingyi Huo, Anusha Devulapally, Hasan Al Maruf, Krishnakumar Nair, Meena Arunachalam, Gulsum Gudukbay Akbulut, Mahmut T. Kandemir, Narayanan Vijaykrishnan
MICRO6
2023 Optimizing CPU Performance for Recommendation Systems At-Scale
abstract
Deep Learning Recommendation Models (DLRMs) are very popular in personalized recommendation systems and are a major contributor to the data-center AI cycles. Due to the high computational and memory bandwidth needs of DLRMs, specifically the embedding stage in DLRM inferences, both CPUs and GPUs are used for hosting such workloads. This is primarily because of the heavy irregular memory accesses in the embedding stage of computation that leads to significant stalls in the CPU pipeline. As the model and parameter sizes keep increasing with newer recommendation models, the computational dominance of the embedding stage also grows, thereby, bringing into question the suitability of CPUs for inference. In this paper, we first quantify the cause of irregular accesses and their impact on caches and observe that off-chip memory access is the main contributor to high latency. Therefore, we exploit two well-known techniques: (1) Software prefetching, to hide the memory access latency suffered by the demand loads and (2) Overlapping computation and memory accesses, to reduce CPU stalls via hyperthreading to minimize the overall execution time. We evaluate our work on a single-core and 24-core configuration with the latest recommendation models and recently released production traces. Our integrated techniques speed up the inference by up to 1.59x, and on average by 1.4x.
Scott Cheng, Vishwas Kalagi, Vrushabh Sanghavi, Samvit Kaul, Meena Arunachalam, Kiwan Maeng, Adwait Jog, Anand Sivasubramaniam, Mahmut T. Kandemir, Chita R. Das
ISCA6
2019 Architecture-Centric Bottleneck Analysis for Deep Neural Network Applications
abstract
The ever-growing complexity and popularity of machine learning and deep learning applications have motivated an urgent need of effective and efficient support for these applications on contemporary computing systems. In this paper, we thoroughly analyze the various DNN algorithms on three widely used architectures (CPU, GPU, and Xeon Phi). The DNN algorithms we choose for evaluation include i) Unet - for biomedical image segmentation, based on Convolutional Neural Network (CNN), ii) NMT - for neural machine translation based on Recurrent Neural Network (RNN), iii) ResNet-50, and iv) DenseNet - both for image processing based on CNNs. The ultimate goal of this paper is to answer four fundamental questions: i) whether the different DNN networks exhibit similar behavior on a given execution platform? ii) whether, across different platforms, a given DNN network exhibits different behaviors? iii) for the same execution platform and the same DNN network, whether different execution phases have different behaviors? and iv) are the current major general-purpose platforms tuned sufficiently well for different DNN algorithms? Motivated by these questions, we conduct an in-depth investigation of running DNN applications on modern systems. Specifically, we first identify the most time-consuming functions (hotspot functions) across different networks and platforms. Next, we characterize performance bottlenecks and discuss them in detail. Finally, we port selected hotspot functions to a cycle-accurate simulator, and use the results to direct architectural optimizations to better support DNN applications.
Jihyun Ryoo, Mengran Fan, Xulong Tang, Huaipan Jiang, Meena Arunachalam, Sharada Naveen, Mahmut T. Kandemir
HiPC5