Swarnava Dey

dblp:124/6820 · DBLP profile ↗
← Back
17ranked-venue papers
7as first author
10since 2021 · last 2026
0000-0002-3988-1445ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 6 · 1 first-author · 4 since 2021Systems, architecture and hardware · 4 · 2 first-author · 2 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Bias-variance games for tiny model synthesis in resource-constrained Earth Observation systems
Swarnava Dey, Pallab Dasgupta, P. P. Chakrabarti 0001
J. Syst. Archit.1
2025 TinyHAR-NAS: Tuning Lightweight Attention Networks for Human Activity Recognition on the Edge
abstract
Deep Neural Networks (DNNs) have significantly enhanced the baseline performance of Human Activity Recognition (HAR) models by learning patterns directly from raw sensor signals. For wearables, on-device inference is crucial for preserving personally identifiable information and ensuring long battery life, particularly in HAR-based health and wellness applications. However, existing DNN models for HAR are often either too large for wearable devices, designed for simple binary activity classification, or rely on manually crafted features. To address these challenges, we propose a scalable solution comprising a tunable tiny attention condenser-based architecture and a Deep Q-Learning-based tuner. Together, these components enable the generation of compact, wearable-friendly DNN models for complex activity recognition tasks. Experimental results demonstrate that the proposed methodology achieves state-of-the-art accuracy on the HAR-Box and UCI-HAR datasets, with model sizes under 1 MB, making it suitable for resource-constrained devices.
Urmi Jana, Shalini Mukhopadhyay, Swarnava Dey, Arijit Mukherjee, Arpan Pal 0001
IJCNN3
2025 Tuning Distillation to Generate Edge-friendly All-rounder Models
abstract
Deep learning models for edge deployments must be small, efficient, and robust. Knowledge distillation from large foundation models (FMs) can help, as they capture rich, transferable representations from large, multimodal datasets. However, two key challenges hinder this process: (1) extreme model compression based on a test dataset often leads to brittleness under distribution shifts, and (2) the distribution gap between an FM’s training data and a small model’s target dataset makes standard distillation methods ineffective.We propose a tunable dynamic loss curriculum for knowledge distillation to address these issues. Experiments show that small encoders trained with our approach achieve balanced transfer learning performance across both primary and out-of-distribution tasks. For instance, a tiny GPT-like model effectively transfers to sentiment classification and language modeling. Likewise, a tiny ResNet trained on CIFAR-10 achieves 10% higher accuracy on corrupted CIFAR-10 than state-of-the-art baselines for tiny models. While it trails dedicated robustness training by 4%, our method ensures superior adaptability across diverse datasets and tasks.
Swarnava Dey, Arijit Mukherjee, Arpan Pal 0001
SMC1
2024 Demo: Stress Detection on Tiny Edge Device with GSR Sensor
abstract
Stress management is paramount to maintaining optimal health and well-being; stress builds up in spikes, causing problems like hypertension and anxiety, necessitating personalized interventions delivered in real-time through wearable technology. This work underscores the pivotal role of unobtrusive stress detection and presents the development of auto-generated compact models tailored for on-device inference through Neural Architecture Search (NAS). These models aim to facilitate efficient stress monitoring directly on low-power devices, representing a promising avenue for advancing personalized healthcare with continuous monitoring and effective digital interventions.
Shalini Mukhopadhyay, Varsha Sharma, Dibyanshu Jaiswal, Swarnava Dey, Avik Ghose
MobiSys4
2023 DietCNN: Multiplication-free Inference for Quantized CNNs
abstract
The rising demand for networked embedded systems with machine intelligence has been a catalyst for sustained attempts by the research community to implement Convolutional Neural Networks (CNN) based inferencing on embedded resource-limited devices. Redesigning a CNN by removing costly multiplication operations has already shown promising results in terms of reducing inference energy usage. This paper proposes a new method for replacing multiplications in a CNN by table look-ups. Unlike existing methods that completely modify the CNN operations, the proposed methodology preserves the semantics of the major CNN operations. Conforming to the existing mechanism of the CNN layer operations ensures that the reliability of a standard CNN is preserved. It is shown that the proposed multiplication-free CNN, based on a single activation codebook, can achieve 4.7x, 5.6x, and 3.5x reduction in energy per inference in an FPGA implementation of MNIST-LeNet-5, CIFAR10-VGG-11, and Tiny ImageNet-ResNet-18 respectively. Our results show that the DietCNN approach significantly improves the resource consumption and latency of deep inference for smaller models, often used in embedded systems. Our code is available at: https://github.com/swadeykgp/DietCNN
Swarnava Dey, Pallab Dasgupta, P. P. Chakrabarti 0001
IJCNN1
2023 Challenges of Accurate and Efficient AutoML
abstract
Embedded Artificial Intelligence (AI) is becoming increasingly important in the field of healthcare where such AI enabled devices are utilized to assist physicians, clinicians, and surgeons in their diagnosis, rehabilitation and therapy planning. However, it is still a challenging task to come up with an accurate and efficient machine learning model for resource-limited devices that work$24\times 7$.. It requires both intuition and experience. This dependence on human expertise and reliance on trial-and-error-based design methods create impediments to the standard processes of effort estimation, design phase planning, and generating service-level agreements for projects that involve AI-enabled MedTech devices. In this paper, we present AutoML search from an algorithmic perspective, instead of a more prevalent optimization or black-box tool view. We briefly present and point to case studies that demonstrate the efficacy of the automation approach in terms of productivity improvements. We believe that our proposed method can make AutoML more amenable to the applications of software engineering principles and also accelerate biomedical device engineering, where there is a high dependence on skilled human resources.
Swarnava Dey, Avik Ghose, Soumik Das
ASE1
2023 Demo: On-device Puff Detection System for Smoking Cessation
abstract
Customized, on-device applications that provide timely interventions about smoking episodes are very helpful for smoking cessation. For this, real-time detection of smoking puffs are necessary through unobtrusive wearable devices. This work demonstrates auto-generated tiny puff detection models for on-device inference on low-power wearable devices.
Shalini Mukhopadhyay, Swarnava Dey, Avik Ghose
MobiSys2
2023 Demo Abstract: Lightweight Attention Network for Time Series Classification on Edge
abstract
In this work, we present a lightweight attention network to perform Time Series Classification on Edge devices. We evaluate the merit of our system on a Human Activity Recognition dataset and show the demonstration with the help of a Wearable device (Smartwatch) with IMU sensors.
Shalini Mukhopadhyay, Swarnava Dey, Arpan Pal 0001, Ashwin S
SenSys2
2022 Automated Generation of Tiny Model for Real-Time ECG Classification on Tiny Edge Devices
abstract
Continuous monitoring of cardiac health through single-lead wearable Electrocardiogram (ECG), is important for paroxysmal Atrial Fibrillation (AF) detection. Wearable ECG straps, watches, and implantable loop recorders (ILR) are based on this paradigm. These devices are used by medical professionals to view data from multiple patients, perform continuous monitoring and analysis to provide immediate care to patients. These monitoring devices display simple health screening alerts to the subjects and generate distress signals for people working outdoors or in isolated environments with intermittent Internet connectivity. Hence, low-memory, low-power, low-latency on-device inference becomes very important. This work aims at realizing such solutions by providing a framework to generate tiny (less than 256 KB) Deep Neural Networks customized for typical microcontrollers (MCU) used in those devices.
Shalini Mukhopadhyay, Swarnava Dey, Avik Ghose, Aakash Tyagi
SenSys2
2022 Accelerated Fire Detection and Localization at Edge
abstract
Fire-related incidents continue to be reported as a leading cause of life and property destruction. Automated fire detection and localization (AFDL) systems have grown in importance with the evolution of applied robotics, especially because use of robots in disaster situations can lead to avoidance of human fatality. The importance of AFDL on resource-constrained devices has further grown, as most unmanned vehicles (drones or ground vehicles) are battery operated with limited computational capacity, the disaster situations cannot guarantee uninterrupted communication with high-end resources in the cloud, and yet faster response time is a prime necessity. Traditional computer vision–based techniques require hand-engineered features on a case-by-case basis. Deep Learning –based classifiers perform well for fire/no-fire classification due to the availability of large datasets for training; however, a dearth of good fire localization datasets renders the localization performance below par. We have tried to address both problems with a multi-task learned cascaded model that triggers localization workflow only if the presence of fire is detected, through a strong classifier trained on available large fire datasets. This presents only fire images to a relatively weaker localization model, reducing false positives, false negatives, and thereby improving overall AFDL accuracy. The multi-task learning (MTL) approach for end-to-end training of a stitched classifier and object localizer model on diverse datasets enabled us to build a strong fire classifier and feature extractor. It also resulted in a single unified model, capable of running on “on-board” compute infrastructure without compromising on accuracy. To achieve the target inference rate for the AFDL deployment, we have investigated the effect of quantization and compression due to hardware acceleration on an MTL model. This article presents an approach to automate the hardware-software co-design to find the optimum parameter partitioning for a given MTL problem, especially when some parts of the model are hardware accelerated. We present combined evaluation results showing that our methodology and the corresponding AFDL model strikes a balance between the frames inferred per second and several accuracy metrics. We report fire localization accuracy in terms of mean average precision (object detection), which has not been done earlier for embedded AFDL systems.
Arijit Mukherjee, Jayeeta Mondal, Swarnava Dey
ACM Trans. Embed. Comput. Syst.3
2019 Edge Acceleration of Deep Neural Networks
abstract
Running deep learning algorithms at the edge is a necessity in many industrial use-cases, especially in applications that use robots and drones in disaster recovery, surveillance, oil & gas operations etc. Current state of the art deep learning algorithms are extremely efficient in analysing image, audio, video and other time-series signals. However, their performance degrades considerably on constrained edge devices. In this demo, we show how standard pretrained CNN (Convolutional Neural Network) models can be partitioned for efficient parallel execution between constrained devices and also achieve real-time response.
Jayeeta Mondal, Swarnava Dey, Arijit Mukherjee, Jeet Dutta, Arpan Pal 0001, Balamurali P
MobiSys2
2018 (WIP) At Most M - A Flexible Redundancy Model for Cloud Robotics
abstract
The current trend in industry to augment lightweight and inexpensive robots with cloud-based distributed intelligence for executing complex, collaborative tasks has led to the emergence of fields like Internet of Robotic Things (IoRT) and Cloud Robotics (CR). However, reliable execution of orchestrated jobs in a networked setup with chances of infrastructural failures remains a concern. The only perceivable solution is resource redundancy, where the traditional approaches are not cost effective for tightly budgeted IoRT/CR deployments. In this work, an agile redundancy model, At most M, is proposed, which provides a handle to tune the trade-off between resource cost and reliability. The model is evaluated in a simulated warehouse environment with robots, drones, AGVs and private cloud servers deployed to accomplish multiple pickup and delivery tasks. The benefits w.r.t. resource usage cost by using our optimized redundancy model is illustrated and the trade-off between cost and reliability is demonstrated.
Swagata Biswas, Swarnava Dey, Arijit Mukherjee
IEEE CLOUD2
2014 Dual scheme phone: a user assisted mechanism to effectively run sensor analytics applications on smartphones
abstract
In recent times, many human-centric applications are being developed to leverage the diverse range of sensors present in Android smartphones. Smartphones are also being used as edge network gateways for fusing data from multiple sensors. These applications often run as background services and usuall
Swarnava Dey, Arijit Mukherjee, Pubali Datta, Himadri Sekhar Paul, Anupam Basu
MobiQuitous1
2014 Facilitating continued run of sensor data analytics services using user driven proactive memory reclamation scheme
abstract
Smartphones are currently being used to develop diverse range of applications (apps) involving sensors. These apps generally acquire and analyze sensor data and are usually implemented as background services. The importance values of Android processes are in a hierarchy of foreground, visible, background etc. in decreasing order of importance. Whenever a new process arrives, it may necessitate removal of old and less important processes for reclaiming memory. Current smartphones do not provide any options through which user's idea of priority can override that of the system defaults. In this work we present an implementation that enables the user to obtain alerts on system load and recommendations to proactively kill a set of processes to reclaim system memory. This enables user selected background process to be spared from the standard android policy of process termination, in lieu of foreground apps, relatively unimportant from user perspective, during that period. We show that manual reclaiming of memory based on recommendations from our app, reduces the automatic killing and measurement lag experienced by a sensor analytics app under test. This work is redundant if processing power and main memory of a smartphone is always surplus than required for its normal usage.
Swarnava Dey, Pubali Datta, Arijit Mukherjee, Himadri Sekhar Paul, Anupam Basu
SenSys1
2013 Challenges of Using Edge Devices in IoT Computation Grids
abstract
Internet of Things (IoT) has the potential to become a technology revolution with a vision of creating very large scale network, comprising of unprecedented number of connected devices. These devices, often referred to as smart items or intelligent things can be home appliances, healthcare devices, vehicles, buildings, factories and almost anything networked and fitted with sensors, actuators, embedded computers. There has been sustained research work and standardization effort from different IoT perspectives like integration of sensor and RFID devices to the Internet. With the increasing trend of gathering business insights from unstructured data, the high volume of data generated by such devices is also of interest. Cloud based data mining platforms are suitable for analyses of such data and researchers have proposed architectures where personal mobile phones can act as Edge Gateway between the sensor network and cloud analytics platform. It seems that the surge in the volume of data generated by huge number of Smart Items can only be matched if a large percentage of mobile users start sharing the computation capability of their personal devices and work together towards true Participatory Computing in the IoT systems. In this work we try to understand the challenges associated with running computation jobs on the mobile devices using different types of workload often observed in IoT applications. Based on the insights gained from experiments performed by us, we propose a scheme where mobile phones, residential gateways and other edge devices offer free slots to servers in a cloud based data analytics system. Based on the free time slots offered by the mobile phones, if commensurately sized computational jobs can be scheduled, the unpredictability associated with using mobile phones as grid resources can be solved.
Swarnava Dey, Arijit Mukherjee, Himadri Sekhar Paul, Arpan Pal 0001
ICPADS1
2013 Utilising condor for data parallel analytics in an IoT context - An experience report
abstract
The current emphasis on sensor-based intelligent and ubiquitous systems, more commonly known as “cyber-physical systems”, has the potential to give rise to a new generation of systems and services encompassing several domains such as e-Governance, healthcare, transportation, waste management, energy & utilities, insurance, etc., resulting in the metamorphosis of the Internet as we see it, into the Internet of Things (IoT). One probable commonality in each of these services will be the abundance of different types of data from different sources with the success of the systems depending on real-time or near real-time analysis of data. Such analyses are normally performed via well-known algorithms with a time-constraint on the execution, thus creating a requirement for parallel execution techniques. Some of these analyses may have a higher frequency of execution on a relatively small set of data, in which case the current big-data frameworks may actually add an overhead. Further, the frameworks like Hadoop demand the algorithms to be mapped onto a particular paradigm, which may not always be a suitable option. This paper, which is a work-in-progress, provides an experience report on the use of Condor, a well known Grid framework, for data-parallel “black-box” style execution of analysis algorithms in the context of Internet of Things. We concentrate on algorithms which are already in use, and can be partitioned into data-parallel subtasks without any modification and use Condor, which has traditionally been used for high-performance or high-throughput computing, as the execution framework.
Arijit Mukherjee, Swarnava Dey, Himadri Sekhar Paul, Batsayan Das
WiMob2
2012 An Erasure Coded Archival Storage System
abstract
There is an ever increasing need of storage capacity for storage of digital archives and historical data-digital preservation, because of regulatory and compliance requirements. There is an increasing interest in disk based archival system. Major technical challenges in creating large disk based storage archive are - providing large capacity at low costs, large read and write throughput, data integrity and sustaining hardware and operating system refresh. In this paper we present the architecture and working principle of an archival storage system that uses an erasure-coded redundancy scheme. We present the design of a Quality of Service (QoS) framework that tries to achieve an optimum balance between file availability, performance and system availability. The design includes a file encoding and placement scheme that allows files to be read from the archive without the need to access any metadata. Finally, we present the results obtained from running an experimental setup on Amazon Web Services.
Prateep Misra, Nilanjan Roy, Soumitra Naskar, Swarnava Dey
ICPADS4