Avijoy Chakma

dblp:214/9679 · DBLP profile ↗
← Back
13ranked-venue papers
4as first author
11since 2021 · last 2026
0000-0001-6818-8327ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 1 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 4 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
YearPublicationVenuePosition
2026 Re-educating Educated Ones: A Case Study on Chakma Language Revitalization in Chittagong Hill Tracts
abstract
Indigenous languages face significant cultural oppression from official state languages, particularly in the Global South. We investigate the Bangladeshi Chakma language revitalization movement, a community grappling with language liquidity and amalgamation into the dominant Bengali language. Our six-month-long qualitative study involving interviews and focus group discussions with Chakma language learning stakeholders uncovered existing community socio-economic challenges and resilience strategies. We noted the need for culturally grounded digital tools and resources. We propose an ICT-mediated community-centric framework for Indigenous language revitalization in the Global South, emphasizing the integration of historical identity elements, stakeholder-defined requirements, and effective digital engagement strategies to empower communities in preserving their linguistic and cultural heritage.
Avijoy Chakma, Adity Khisa, Soham Khisa, Jannatun Noor 0001, Sharifa Sultana
CHI1
2025 The Good, The Bad and The Ugly: The Opportunities, Challenges and the Mitigation Strategies of the Young Indigenous Social Media Users of the Chittagong Hill Tracts in Bangladesh
Ishmam Bin Rofi, Mashiyat Mahjabin Eshita, Avijoy Chakma, Md. Sabbir Ahmed 0002, S. M. Taiabul Haque, Jannatun Noor 0001
CHI3
2025 VaCDA: Variational Contrastive Alignment-based Scalable Human Activity Recognition
abstract
Technological advancements have led to the rise of wearable devices with sensors that continuously monitor user activities, generating vast amounts of unlabeled data. This data is challenging to interpret, and manual annotation is labor-intensive and error-prone. Additionally, data distribution is often heterogeneous due to device placement, type, and user behavior variations. As a result, traditional transfer learning methods perform sub-optimally, making it difficult to recognize daily activities. To address these challenges, we deploy a variational autoencoder (VAE) to learn a shared, low-dimensional latent space from available sensor data. This space generalizes data across diverse sensors, mitigating heterogeneity and aiding robust adaptation to the unlabeled data. We integrate contrastive learning to enhance feature representation by aligning instances of the same class across domains while separating different classes. Eventually, we propose a novel Variational Contrastive Domain Adaptation (VaCDA), a multi-source domain adaptation framework combining VAEs and contrastive learning to improve feature representation and reduce heterogeneity between source and target domains. We evaluate VaCDA on multiple publicly available datasets across three heterogeneity scenarios: cross-person, cross-position, and cross-device. VaCDA significantly outperforms the baselines in cross-position and cross-device scenarios.
Soham Khisa, Avijoy Chakma
ICMLA2
2024 AcouDL: Context-Aware Daily Activity Recognition from Natural Acoustic Signals
abstract
The ubiquitousness of smart and wearable devices with integrated acoustic sensors in modern human lives presents tremendous opportunities for recognizing human activities in our living spaces through ML-driven applications. However, their adoption is often hindered by the requirement of large amounts of labeled data during the model training phase. Integration of contextual metadata has the potential to alleviate this since the nature of these meta-data is often less dynamic (e.g. cleaning dishes, and cooking both can happen in the kitchen context) and can often be annotated in a less tedious manner (a sensor always placed in the kitchen). However, most models do not have good provisions for the integration of such meta-data information. Often, the additional metadata is leveraged in the form of multi-task learning with sub-optimal outcomes. On the other hand, reliably recognizing distinct in-home activities with similar acoustic patterns (e.g. chopping, hammering, knife sharpening) poses another set of challenges. To mitigate these challenges, we first show in our preliminary study that the room acoustics properties such as reverberation, room materials, and background noise leave a discernible fingerprint in the audio samples to recognize the room context and proposed AcouDL as a unified framework to exploit room context information to improve activity recognition performance. Our proposed self-supervision-based approach first learns the context features of the activities by leveraging a large amount of unlabeled data using a contrastive learning mechanism and then incorporates this feature induced with a novel attention mechanism into the activity classification pipeline to improve the activity recognition performance. Extensive evaluation of AcouDL on three datasets containing a wide range of activities shows that such an efficient feature fusion-mechanism enables the incorporation of metadata that helps to better recognition of the activities under challenging classification scenarios with 0.7-3.5% macro F1 score improvement over the baselines.
Avijoy Chakma, Anirban Das 0005, Abu Zaher Md Faridee, Suchetana Chakraborty, Sandip Chakraborty 0001, Nirmalya Roy
SMARTCOMP1
2023 A Systematic Study on Object Recognition Using Millimeter-wave Radar
abstract
Millimeter-wave (MMW) radar is becoming an essential sensing technology in smart environments due to its light and weather-independent sensing capability. Such capabilities have been widely explored and integrated with intelligent vehicle systems, often deployed in industry-grade MMW radars. However, industry-grade MMW radars are often expensive and difficult to attain for deployable community-purpose smart environment applications. On the other hand, commercially available MMW radars pose hidden underpinning challenges that are yet to be well investigated for tasks such as recognizing objects, and activities, real-time person tracking, object localization, etc. Such tasks are frequently accompanied by image and video data, which are relatively easy for an individual to obtain, interpret, and annotate. However, image and video data are light and weather-dependent, vulnerable to the occlusion effect, and inherently raise privacy concerns for individuals. It is crucial to investigate the performance of an alternative sensing mechanism where commercially available MMW radars can be a viable alternative to eradicate the dependencies and preserve privacy issues. Before championing MMW radar, several questions need to be answered regarding MMW radar’s practical feasibility and performance under different operating environments. To answer the concerns, we have collected a dataset using commercially available MMW radar, Automotive mmWave Radar (AWR2944) from Texas Instruments, and reported the optimum experimental settings for object recognition performance using several deep learning algorithms in this study. Moreover, our robust data collection procedure allows us to systematically study and identify potential challenges in the object recognition task under a cross-ambience scenario. We have explored the potential approaches to overcome the underlying challenges and reported extensive experimental results.
Maloy Kumar Devnath, Avijoy Chakma, Mohammad Saeid Anwar, Emon Dey, Zahid Hasan 0001, Marc Conn, Biplab Pal, Nirmalya Roy
SMARTCOMP2
2023 HeteroSys: Heterogeneous and Collaborative Sensing in the Wild
abstract
Advances in Internet-of-Things, artificial intelligence, and ubiquitous computing technologies have contributed to building the next generation of context-aware heterogeneous systems with robust interoperability to control and monitor the environmental variables of smart environments. Motivated by this, we propose HeteroSys, an end-to-end multi-functional smart IoT-based system prototype for heterogeneous and collaborative sensing in a smart IoT-based environment. A unique characteristic of HeteroSys is that it relies on Home Assistant (HA) to collate heterogeneous sensors (e.g., passive infrared sensors (PIR), reed (door) switches, object tags, wearable wrist-mounted, water leak sensors, and internet protocol cameras), and uses a variety of networking protocols such as Zigbee open standard for mesh networking, WiFi, and Bluetooth Low Energy (BLE) for communication. The reliance on HA (and its broad community support) makes HeteroSys ideal for various applications such as object detection, human activity recognition and behavior patterns. We articulated the development phase, integration, testing challenges and evaluation of the HeteroSys. We conducted an extensive 24-hour longitudinal data collection from 5 participants performing 6 activities by deploying in an indoor home environment. Our assessment of the acquired dataset reveals that the representations learned using deep learning architecture aid in improving the detection of activities to 83.1% accuracy.
Indrajeet Ghosh, Adam Goldstein, Avijoy Chakma, Jade Freeman, Timothy Gregory, Niranjan Suri, Sreenivasan Ramasamy Ramamurthy, Nirmalya Roy
SMARTCOMP3
2022 Semi-supervised Multi-source Domain Adaptation in Wearable Activity Recognition
abstract
The scarcity of labeled data has traditionally been the primary hindrance in building scalable supervised deep learning models that can retain adequate performance in the presence of various heterogeneities in sample distributions. Domain adaptation tries to address this issue by adapting features learned from a smaller set of labeled samples to that of the incoming unlabeled samples. The traditional domain adaptation approaches normally consider only a single source of labeled samples, but in real world use cases, labeled samples can originate from multiple-sources – providing motivation for multi-source domain adaptation (MSDA). Several MSDA approaches have been investigated for wearable sensor-based human activity recognition (HAR) in recent times, but their performance improvement compared to single source counterpart remained marginal. To remedy this performance gap that, we explore multiple avenues to align the conditional distributions in addition to the usual alignment of marginal ones. In our investigation, we extend an existing multi-source domain adaptation approach under semi-supervised settings. We assume the availability of partially labeled target domain data and further explore the pseudo labeling usage with a goal to achieve a performance similar to the former. In our experiments on three publicly available datasets, we find that a limited labeled target domain data and pseudo label data boost the performance over the unsupervised approach by 10-35% and 2-6%, respectively, in various domain adaptation scenarios.
Avijoy Chakma, Abu Zaher Md Faridee, Raghuveer M. Rao, Nirmalya Roy
DCOSS1
2022 PerMTL: A Multi-Task Learning Framework for Skilled Human Performance Assessment
abstract
Intelligent and complex human motion analysis can help design the next generation IoT and AR/VR systems for automated human performance assessment. Such an automated system can help advocate the interpretability and translatability of complex human motions, intelligent motion feedback, and fine-grained motion skill assessment to design next-generation interactive human-machine teaming systems. Motivated by this, we design a wearable sensing framework for assessing the players’ performance and consider a live badminton game as our use case. Generally, the players on the field try to improve their performance by focusing on fast and synchronous coordination of their limbs’ reflex actions to have the ideal body postures to perform the desired shot. Learning the minute dissimilarities and distinctive traits from each limb of the players simultaneously can help assess the players’ performance and specific skillsets during a game. This paper proposes a multi-task learning framework, PerMTL to learn the shared features from each player’s limb. The PerMTL comprises a task-specific regressor output layer that helps to determine the dissimilarities and distinctive traits between the player’s limbs for collective inference in a body sensor network (BSN) environment. We evaluate the PerMTL framework using publicly available Badminton Activity Recognition (BAR) and Daily and Sports Activities (DSA) datasets. Empirical results indicate that PerMTL achieves R2Score of ≈ 82% in predicting the players’ performance.
Indrajeet Ghosh, Avijoy Chakma, Sreenivasan Ramasamy Ramamurthy, Nirmalya Roy, Nicholas R. Waytowich
ICMLA2
2022 CoDEm: Conditional Domain Embeddings for Scalable Human Activity Recognition
abstract
We explore the effect of auxiliary labels in improving the classification accuracy of wearable sensor-based human activity recognition (HAR) systems, which are primarily trained with the supervision of the activity labels (e.g. running, walking, jumping). Supplemental meta-data are often available during the data collection process such as body positions of the wearable sensors, subjects' demographic information (e.g. gender, age), and the type of wearable used (e.g. smartphone, smart-watch). This information, while not directly related to the activity classification task, can nonetheless provide auxiliary supervision and has the potential to significantly improve the HAR accuracy by providing extra guidance on how to handle the introduced sample heterogeneity from the change in domains (i.e positions, persons, or sensors), especially in the presence of limited activity labels. However, integrating such meta-data information in the classification pipeline is non-trivial - (i) the complex interaction between the activity and domain label space is hard to capture with a simple multi-task and/or adversarial learning setup, (ii) meta-data and activity labels might not be simultaneously available for all collected samples. To address these issues, we propose a novel framework Conditional Domain Embeddings (CoDEm). From the available unlabeled raw samples and their domain meta-data, we first learn a set of domain embeddings using a contrastive learning methodology to handle inter-domain variability and inter-domain similarity. To classify the activities, CoDEm then learns the label embeddings in a contrastive fashion, conditioned on domain embeddings with a novel attention mechanism, enforcing the model to learn the complex domain-activity relationships. We extensively evaluate CoDEm in three benchmark datasets against a number of multi-task and adversarial learning baselines and achieve state-of-the-art nerformance in each avenue.
Abu Zaher Md Faridee, Avijoy Chakma, Zahid Hasan 0001, Nirmalya Roy, Archan Misra
SMARTCOMP2
2022 DeCoach: Deep Learning-based Coaching for Badminton Player Assessment
abstract
Wearable devices have gained immense popularity among various pervasive computing and Internet-of-Things (IoT) applications in the past decade. Sports analytics researchers recently focused on improving a player’s performance to help devise a winning strategy based on the player’s gameplay. Especially in a racquet-based badminton sport, it is assumed that handling the racquet during the gameplay is one of the primary reasons to influence the players’ performance. On the contrary, we posit that the players’ stance, body movements, and posture are equally significant in evaluating a player’s performance during the game. A shot characterized by a recommended posture, stance, and body movements allows a player to play a stroke efficiently, thus aiding the player in guiding the shuttle to strategic spots and making it difficult for the opponent to return the shot and score a point. Relying on this hypothesis, we propose DeCoach, a data-driven framework that leverages the stance and posture of the players and ranks them based on their performances. In this effort, we first employ a deep learning-based algorithm to classify the strokes and stances of the players. Secondly, we propose a distance-based methodology to compare the obtained stance of a player with that of a professional player. Finally, we devise a deep learning-based regressor to predict the player’s performance which commences with ranking based on their performance. We evaluate DeCoach using our in-house dataset, Badminton Activity Recognition (BAR) Dataset that is collected using inertial measurement unit (IMU) sensors by placing them on the upper and lower limbs of the players. The BAR dataset is collected from 11 players in the controlled and uncontrolled environment settings for 12 frequently played shots in the game. Empirical results indicate that DeCoach achieves 89.09% accuracy for strokes detection and R2 score of 88.84% in estimating the players’ performance.
Indrajeet Ghosh, Sreenivasan Ramasamy Ramamurthy, Avijoy Chakma, Nirmalya Roy
Pervasive Mob. Comput.3
2021 A Scalable and Domain Adaptive Respiratory Symptoms Detection Framework using Earables
abstract
The COVID-19 pandemic has brought a devastating impact on human health across the globe, and people are still observing face-masking as a preventive measure to contain the spread of COVID-19. Coughing is one of the major transmission mediums of COVID-19, and early cough detection could play a significant r ole i n p reventing t he s pread o f t his life-threatening virus. Many approaches have been proposed for developing systems to detect coughing and other respiratory symptoms in literature, but earable devices are not well-studied and investigated for respiratory symptom detection. In this work, we posited an acoustic research prototype (earable device) - eSense that has acoustic and IMU sensors embedded into user-convenient earbuds to address the following issues: (i) feasibility of the earables in detecting respiratory symptoms, and (ii) scalability of trained machine learning models in the presence of unseen data samples. We performed experimentation with both shallow and deep learning models on the eSense collected data samples. We observed that the deep learning model outperforms the shallow learning models achieving 97% accuracy. Furthermore, we investigated the scalability of the deep learning model on unseen datasets and noticed that the performance of the deep learning model deteriorates when trained on a particular dataset and tested on an unseen dataset. To mitigate such challenges, we postulated an adversarial domain adaptation technique that helps improve the performance of our respiratory symptoms detection framework by a substantial margin.
Avijoy Chakma, Nirmalya Roy
IEEE BigData2
2020 LASO: Exploiting Locomotive and Acoustic Signatures over the Edge to Annotate IMU Data for Human Activity Recognition
abstract
Annotated IMU sensor data from smart devices and wearables are essential for developing supervised models for fine-grained human activity recognition, albeit generating sufficient annotated data for diverse human activities under different environments is challenging. Existing approaches primarily use human-in-the-loop based techniques, including active learning; however, they are tedious, costly, and time-consuming. Leveraging the availability of acoustic data from embedded microphones over the data collection devices, in this paper, we propose LASO, a multimodal approach for automated data annotation from acoustic and locomotive information. LASO works over the edge device itself, ensuring that only the annotated IMU data is collected, discarding the acoustic data from the device itself, hence preserving the audio-privacy of the user. In the absence of any pre-existing labeling information, such an auto-annotation is challenging as the IMU data needs to be sessionized for different time-scaled activities in a completely unsupervised manner. We use a change-point detection technique while synchronizing the locomotive information from the IMU data with the acoustic data, and then use pre-trained audio-based activity recognition models for labeling the IMU data while handling the acoustic noises. LASO efficiently annotates IMU data, without any explicit human intervention, with a mean accuracy of $0.93$ ($\pm 0.04$) and $0.78$ ($\pm 0.05$) for two different real-life datasets from workshop and kitchen environments, respectively.
Soumyajit Chatterjee, Avijoy Chakma, Aryya Gangopadhyay, Nirmalya Roy, Bivas Mitra, Sandip Chakraborty 0001
ICMI2
2017 Image-based air quality analysis using deep convolutional neural network
abstract
Air pollution may cause many severe diseases. An efficient air quality monitoring system is of great benefit for human health and air pollution control. In this paper, we study image-based air quality analysis, in particular, the concentration estimation of particulate matter with diameters less than 2.5 micrometers (PM2.5). The proposed method uses a deep Convolutional Neural Network (CNN) to classify natural images into different categories based on their PM2.5 concentrations. In order to evaluate the proposed method, we created a dataset that contains total 591 images taken in Beijing with corresponding PM2.5 concentrations. The experimental results demonstrate that our method are valid for image-based PM2.5 concentration estimation.
Avijoy Chakma, Ben Vizena, Tingting Cao, Zhiyuan Jerry Lin
ICIP1