Dhananjay Saikumar

dblp:371/2446 · DBLP profile ↗
← Back
1ranked-venue papers
1as first author
1since 2021 · last 2024
0009-0006-4937-9308ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Efficient and distributed learning · 50% Deep learning architectures and training · 50%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Hardware accelerators and domain-specific architectures · 100%

Topics — the 4 heaviest of 5, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Deep learning architectures and training › neural network training
local learning
0.812024
NeuroFlux: Memory-Efficient CNN Training Using Adaptive Local Learning · EuroSys 2024
Machine learning › Efficient and distributed learning
memory-efficient training
0.812024
NeuroFlux: Memory-Efficient CNN Training Using Adaptive Local Learning · EuroSys 2024
Hardware accelerators and domain-specific architectures
edge accelerator
0.812024
NeuroFlux: Memory-Efficient CNN Training Using Adaptive Local Learning · EuroSys 2024
Hardware accelerators and domain-specific architectures
memory-constrained training
0.812024
NeuroFlux: Memory-Efficient CNN Training Using Adaptive Local Learning · EuroSys 2024

Methods — techniques the papers use, named apart from their topics

auxiliary network · 1.5adaptive batch sizes · 1.5activation caching · 1.5
YearPublicationVenuePosition
2024 NeuroFlux: Memory-Efficient CNN Training Using Adaptive Local Learning
abstract
Efficient on-device Convolutional Neural Network (CNN) training in resource-constrained mobile and edge environments is an open challenge. Backpropagation is the standard approach adopted, but it is GPU memory intensive due to its strong inter-layer dependencies that demand intermediate activations across the entire CNN model to be retained in GPU memory. This necessitates smaller batch sizes to make training possible within the available GPU memory budget, but in turn, results in substantially high and impractical training time. We introduce NeuroFlux, a novel CNN training system tailored for memory-constrained scenarios. We develop two novel opportunities: firstly, adaptive auxiliary networks that employ a variable number of filters to reduce GPU memory usage, and secondly, block-specific adaptive batch sizes, which not only cater to the GPU memory constraints but also accelerate the training process. NeuroFlux segments a CNN into blocks based on GPU memory usage and further attaches an auxiliary network to each layer in these blocks. This disrupts the typical layer dependencies under a new training paradigm - 'adaptive local learning'. Moreover, NeuroFlux adeptly caches intermediate activations, eliminating redundant forward passes over previously trained blocks, further accelerating the training process. The results are twofold when compared to Backpropagation: on various hardware platforms, NeuroFlux demonstrates training speed-ups of 2.3× to 6.1× under stringent GPU memory budgets, and NeuroFlux generates streamlined models that have 10.9× to 29.4× fewer parameters.
Dhananjay Saikumar, Blesson Varghese
EuroSys1