Meenakshi Arunachalam

dblp:39/652 · DBLP profile ↗
← Back
4ranked-venue papers
1as first author
1since 2021 · last 2021
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 4 · 1 first-author · 1 since 2021Systems, architecture and hardware · 2 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
Memory systems · 98% Storage systems · 2% High-performance computing · 0%
Software engineering, system software, and programming languages
1 paper
Compilers and program optimization · 100%

Topics — the 4 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Memory systems › cache
cache performance
0.412019
Co-optimizing memory-level parallelism and cache-level parallelism · PLDI 2019
Memory systems › memory access optimization
memory-level parallelism
0.412019
Co-optimizing memory-level parallelism and cache-level parallelism · PLDI 2019
Storage systems › file systems › distributed file system
parallel file system
0.011995
A Prefetching Prototype for the Parallel File System on the Paragon · SIGMETRICS 1995
Memory systems › cache
prefetching
0.011995
A Prefetching Prototype for the Parallel File System on the Paragon · SIGMETRICS 1995

Methods — techniques the papers use, named apart from their topics

compiler-based optimization · 0.8
YearPublicationVenuePosition
2021 Morphable Convolutional Neural Network for Biomedical Image Segmentation
abstract
We propose a morphable convolution framework, which can be applied to irregularly shaped region of input feature map. This framework reduces the computational footprint of a regular CNN operation in the context of biomedical semantic image segmentation. The traditional CNN based approach has high accuracy, but suffers from high training and inference computation costs, compared to a conventional edge detection based approach. In this work, we combine the concept of morphable convolution with the edge detection algorithms resulting in a hierarchical framework, which first detects the edges and then generate a layer-wise annotation map. The annotation map guides the convolution operation to be run only on a small, useful fraction of pixels in the feature map. We evaluate our framework on three cell tracking datasets and the experimental results indicate that our framework saves ~30% and ~10% execution time on CPU and GPU, respectively, without loss of accuracy, compared to the baseline conventional CNN approaches.
Huaipan Jiang, Anup Sarma, Mengran Fan, Jihyun Ryoo, Meenakshi Arunachalam, Sharada Naveen, Mahmut T. Kandemir
DATE5
2019 Co-optimizing memory-level parallelism and cache-level parallelism
abstract
Minimizing cache misses has been the traditional goal in optimizing cache performance using compiler based techniques. However, continuously increasing dataset sizes combined with large numbers of cache banks and memory banks connected using on-chip networks in emerging manycores/accelerators makes cache hit–miss latency optimization as important as cache miss rate minimization. In this paper, we propose compiler support that optimizes both the latencies of last-level cache (LLC) hits and the latencies of LLC misses. Our approach tries to achieve this goal by improving the parallelism exhibited by LLC hits and LLC misses. More specifically, it tries to maximize both cache-level parallelism (CLP) and memory-level parallelism (MLP). This paper presents different incarnations of our approach, and evaluates them using a set of 12 multithreaded applications. Our results indicate that (i) optimizing MLP first and CLP later brings, on average, 11.31% performance improvement over an approach that already minimizes the number of LLC misses, and (ii) optimizing CLP first and MLP later brings 9.43% performance improvement. In comparison, balancing MLP and CLP brings 17.32% performance improvement on average.
Xulong Tang, Mahmut T. Kandemir, Mustafa Karaköy, Meenakshi Arunachalam
PLDI4
2015 Performance and energy evaluation of data prefetching on intel Xeon Phi
abstract
There is an urgent need to evaluate the existing parallelism and data locality-oriented techniques on emerging manycore machines using multithreaded applications. Data prefetching is a well-known latency hiding technique that comes with various hardware- and software-based implementations in almost all commercial machines. A well-tuned prefetcher can reduce the observed data access latencies significantly by bringing the soonto- be-requested data into the cache ahead of time, eventually improving application execution time. Motivated by this, we present in this paper a detailed performance and power characterization of software (compiler-guided) and hardware data prefetching on an Intel Xeon Phi-based system. Our main contributions are (i) an analysis of the interactions between hardware and software prefetching, showing how hardware prefetching can throttle itself in response to software; (ii) results on the power and energy behavior of prefetching, showing how performance and energy gains outweigh the increased power cost of prefetching; and (iii) an evaluation of the use of intrinsic prefetch instructions to prefetch for applications with difficult-to-detect access patterns.
Diana R. Guttman, Mahmut T. Kandemir, Meenakshi Arunachalam, Vlad Calina
ISPASS3
1995 A Prefetching Prototype for the Parallel File System on the Paragon
abstract
Article Free Access Share on A prefetching prototype for the parallel file systems on the Paragon Authors: Meenakshi Arunachalam School of Computer and Information Science, Syracuse University, Syracuse, NY School of Computer and Information Science, Syracuse University, Syracuse, NYView Profile , Alok Choudhary Syracuse University, Department of Electrical and Computer Engineering, Syracuse, NY Syracuse University, Department of Electrical and Computer Engineering, Syracuse, NYView Profile Authors Info & Claims SIGMETRICS '95/PERFORMANCE '95: Proceedings of the 1995 ACM SIGMETRICS joint international conference on Measurement and modeling of computer systemsMay 1995 Pages 321–322https://doi.org/10.1145/223587.223631Published:01 May 1995Publication History 6citation171DownloadsMetricsTotal Citations6Total Downloads171Last 12 Months6Last 6 weeks0 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteeReaderPDF
Meenakshi Arunachalam, Alok N. Choudhary, Brad Rullman
SIGMETRICS1