Arjun Suresh

dblp:165/2533 · DBLP profile ↗
← Back
9ranked-venue papers
5as first author
5since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 3 first-author · 4 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2026 AXI Marks the Spot: Towards Automated Reverse Engineering in SoC Netlists
abstract
Modern Systems-on-Chip (SoCs) integrate numerous IP blocks interconnected through on-chip buses such as AXI. Gate-level SoC netlists consist of millions of gates and dense interconnect structures, while containing little to no semantic information. The resulting sea-of-gates representation lacks hierarchy, module boundaries, and meaningful signal names, severely complicating structural recovery.
Arjun Suresh, Nils Albartus, Daniel E. Holcomb
ACM Great Lakes Symposium on VLSI1
2025 Security Aware Placement for Split Manufactured Integrated Circuit Design Flows
Arjun Suresh, Hari Narayana Burra, Wei-Huan Chen, Daniel E. Holcomb
ACM Great Lakes Symposium on VLSI1
2025 MLPerf Power: Benchmarking the Energy Efficiency of Machine Learning Systems from μWatts to MWatts for Sustainable AI
abstract
Rapid adoption of machine learning (ML) technologies has led to a surge in power consumption across diverse systems, from tiny IoT devices to massive datacenter clusters. Benchmarking the energy efficiency of these systems is crucial for optimization, but presents novel challenges due to the variety of hardware platforms, workload characteristics, and system-level interactions. This paper introduces MLPerf® Power, a comprehensive benchmarking methodology with capabilities to evaluate the energy efficiency of ML systems at power levels ranging from microwatts to megawatts. Developed by a consortium of industry professionals from more than 20 organizations, coupled with insights from academia, MLPerf Power establishes rules and best practices to ensure comparability across diverse architectures. We use representative workloads from the MLPerf benchmark suite to collect $\mathbf{1, 8 4 1}$ reproducible measurements from 60 systems across the entire range of ML deployment scales. Our analysis reveals trade-offs between performance, complexity, and energy efficiency across this wide range of systems, providing actionable insights for designing optimized ML solutions from the smallest edge devices to the largest cloud infrastructures. This work emphasizes the importance of energy efficiency as a key metric in the evaluation and comparison of the ML system, laying the foundation for future research in this critical area. We discuss the implications for developing sustainable AI solutions and standardizing energy efficiency benchmarking for ML systems.
Arya Tschand, Arun Tejusve Raghunath Rajan, Sachin Idgunji, Jeremy Holleman, Csaba Király 0002, Pawan Ambalkar, Ritika Borkar, Ramesh Chukka, Trevor Cockrell, Oliver Curtis, Grigori Fursin, Miro Hodak, Hiwot Kassa, Anton Lokhmotov, Dejan Miskovic, Yuechao Pan, Manu Prasad Manmathan, Liz Raymond, Tom St. John, Arjun Suresh, Rowan Taubitz, Sean Zhan, Scott Wasson, David Kanter, Vijay Janapa Reddi
HPCA21
2025 ODYSSEE: Oyster Detection Yielded by Sensor Systems on Edge Electronics
abstract
Oysters are a vital keystone species in coastal ecosystems, providing significant economic, environmental, and cultural benefits. As the importance of oysters grows, so does the relevance of autonomous systems for their detection and monitoring. However, current monitoring strategies often rely on destructive methods. While manual identification of oysters from video footage is non-destructive, it is time-consuming, requires expert input, and is further complicated by the challenges of the underwater environment. To address these challenges, we propose a novel pipeline using stable diffusion to augment a collected real dataset with photorealistic synthetic data. This method enhances the dataset used to train a YOLOv10-based vision model. The model is then deployed and tested on an edge platform; Aqua2, an Autonomous Underwater Vehicle (AUV), achieving a state-of-the-art 0.657 mAP@50 for oyster detection.
Xiaomin Lin 0002, Vivek Mange, Arjun Suresh, Bernhard Neuberger, Aadi Palnitkar, Brendan Campbell, Kleio Baxevani, Jeremy Mallette, Alhim Vera, Markus Vincze, Ioannis M. Rekleitis, Herbert G. Tanner, Yiannis Aloimonos
ICRA3
2024 Croissant: A Metadata Format for ML-Ready Datasets
abstract
Data is a critical resource for machine learning (ML), yet working with data remains a key friction point. This paper introduces Croissant, a metadata format for datasets that creates a shared representation across ML tools, frameworks, and platforms. Croissant makes datasets more discoverable, portable, and interoperable, thereby addressing significant challenges in ML data management. Croissant is already supported by several popular dataset repositories, spanning hundreds of thousands of datasets, enabling easy loading into the most commonly-used ML frameworks, regardless of where the data is stored. Our initial evaluation by human raters shows that Croissant metadata is readable, understandable, complete, yet concise.
Mubashara Akhtar, Omar Benjelloun, Costanza Conforti, Luca Foschini 0002, Joan Giner-Miguelez, Pieter Gijsbers, Sujata S. Goswami, Nitisha Jain, Michalis Karamousadakis, Michael Kuchnik, Satyapriya Krishna, Sylvain Lesage, Quentin Lhoest, Pierre Marcenac, Manil Maskey, Peter Mattson, Luis Oala, Hamidah Oderinwale, Pierre Ruyssen, Tim Santos, Rajat Shinde, Elena Simperl, Arjun Suresh, Goeffry Thomas, Slava Tykhonov, Joaquin Vanschoren, Susheel Varma, Jos van der Velde, Steffen Vogler, Carole-Jean Wu
NeurIPS23
2020 Hardware Accelerators for Edge Enabled Machine Learning
abstract
The proliferation of IoT devices in recent years has resulted in an exponential increase in data being transmitted over the internet. The traffic is slated for further increase in the coming years and will result in excessive network congestion and high latency. To alleviate this problem, an alternate approach needs to be considered. A prominent option would be to move the computing domain to the edge device. This option is constrained due to reduced computing, storage and power available on the edge. A novel approach combining both software and hardware solutions is required to perform analytics at the edge. This paper proposes an architecture for analysing data on the edge, combining hardware and software solutions. The proposed methodology explores machine learning algorithms for edge computing combined with the use of hardware accelerators to achieve truly intelligent edge devices. A qualitative and quantitative comparison of performance of various algorithms on CPU, GPU, FPGA platforms is carried out. A machine learning model for predicting Remaining Useful Life (RUL) for a multivariate time series dataset is developed and its deployment on the edge is discussed. The results of the experiments carried out are promising and hold potential for further research.
Arjun Suresh, Bhargava N. Reddy, Ch. Renu Madhavi
TENCON1
2018 TTLG - An Efficient Tensor Transposition Library for GPUs
abstract
This paper presents a Tensor Transposition Library for GPUs (TTLG). A distinguishing feature of TTLG is that it also includes a performance prediction model, which can be used by higher level optimizers that use tensor transposition. For example, tensor contractions are often implemented by using the TTGT (Transpose-Transpose-GEMM-Transpose) approach - transpose input tensors to a suitable layout and then use high-performance matrix multiplication followed by transposition of the result. The performance model is also used internally by TTLG for choosing among alternative kernels and/or slicing/blocking parameters for the transposition. TTLG is compared with current state-of-the-art alternatives for GPUs. Comparable or better transposition times for the "repeated-use" scenario and considerably better "single-use" performance are observed.
Jyothi Vedurada, Arjun Suresh, Aravind Sukumaran-Rajam, Changwan Hong, Ajay Panyala, Sriram Krishnamoorthy, V. Krishna Nandivada, Rohit Kumar Srivastava, P. Sadayappan
IPDPS2
2017 Compile-time function memoization
Arjun Suresh, Erven Rohou, André Seznec
CC1
2015 Intercepting Functions for Memoization: A Case Study Using Transcendental Functions
abstract
Memoization is the technique of saving the results of executions so that future executions can be omitted when the input set repeats. Memoization has been proposed in previous literature at the instruction, basic block, and function levels using hardware, as well as pure software--level approaches including changes to programming language. In this article, we focus on software memoization for procedural languages such as C and Fortran at the granularity of a function. We propose a simple linker-based technique for enabling software memoization of any dynamically linked pure function by function interception and illustrate our framework using a set of computationally expensive pure functions—the transcendental functions. Transcendental functions are those that cannot be expressed in terms of a finite sequence of algebraic operations (trigonometric functions, exponential functions, etc.) and hence are computationally expensive. Our technique does not need the availability of source code and thus can even be applied to commercial applications, as well as applications with legacy codes. As far as users are concerned, enabling memoization is as simple as setting an environment variable. Our framework does not make any specific assumptions about the underlying architecture or compiler toolchains and can work with a variety of current architectures. We present experimental results for a x86-64 platform using both gcc and icc compiler toolchains, and an ARM Cortex-A9 platform using gcc. Our experiments include a mix of real-world programs and standard benchmark suites: SPEC and Splash2x. On standard benchmark applications that extensively call the transcendental functions, we report memoization benefits of up to 50% on Intel Ivy Bridge and up to 10% on ARM Cortex-A9. Memoization was able to regain a performance loss of 76% in bwaves due to a known performance bug in the GNU implementation of the pow function. The same benchmark on ARM Cortex-A9 benefited by more than 200%.
Arjun Suresh, Bharath Narasimha Swamy, Erven Rohou, André Seznec
ACM Trans. Archit. Code Optim.1