Subhajit Chakrabarty

dblp:214/8066 · DBLP profile ↗
← Back
9ranked-venue papers
5as first author
5since 2021 · last 2025
0000-0003-0818-3190ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 5 · 2 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 4 · 3 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
YearPublicationVenuePosition
2025 Geospatial Analysis Using Transformer on TOAR and Meteorological Data for Early Warning of Particulate Matter Exceedance
abstract
Forecasting fine particulate matter (PM 2.5) accurately and transparently is essential for public health and environmental policy. The aim of this work is to build an environmental model for early warning of particulate matter exceedance of regulatory limits in the United States. This study used a transformer and TOAR based deep learning model to predict daily PM 2.5 concentrations up to seven days in advance using Top-of-Atmosphere Reflectance (TOAR) and meteorological variables as inputs. Unlike many existing approaches that rely on aerosol optical depth (AOD), the model leverages TOAR directly, avoiding intermediate atmospheric retrieval steps that can introduce bias or uncertainty. The model is trained and evaluated on monitoring stations across the United States using EPA ground-truth data. The key contribution of this work is to build on prior TOAR based transformer forecasting by integrating attention weights for modeling exceedance together with gradient based saliency to interpret model sensitivity in this domain. The second key contribution is identification of spatial differences in lagged behavior for persistence of PM 2.5. The model had an overall$\mathbf{R}$-square of$\mathbf{0. 9 4 5 4}$. The average attention weight, for cases in which PM 2.5 exceedances were predicted with at least 3 days of lead time, showed high importance of lagged variables. TOAR Band 1, temperature and dew point consistently received the highest weights indicating the importance of aerosol presence, particle formation and transport. Gradient-based saliency showed highest values between context indices (days) of 15 and 22. Regional (spatial) analysis showed that the northern region had the highest memory, and the southern region had the shortest memory. Exceedance forecasting showed that the most correct alerts were forecasted with 3 to 4 days of lead time. In summary, this work may provide a stronger foundation for building interpretable, accurate, and lead-time aware pollution forecasting systems that support environmental monitoring and proactive intervention.
Mridula Mavuri, Subhajit Chakrabarty, Udaysinh Rathod, Devesh Sarda
BIBM2
2024 Synthetic Generation of Intermediate MR Images of Lumbar Spine Stenosis for Segmentation
abstract
Lumbar spinal stenosis is a narrowing or constriction in the lumbar spine. Magnetic Resonance Imaging (MRI) is preferred for its diagnosis. Limited research has been conducted with the intent of making synthetic and pseudo in-painting datasets. The objective of this study was to develop a model to create a slice of MR image in the subsequent spinal disc, given an MR image slice from previous disc. To the best of our knowledge, this study is the first to explore the effectiveness of creating pseudo in-painting in the datasets for Lumbar Spine Stenosis. We selected a publicly available dataset, named Lumbar Spine MR images Dataset, for reproducibility. We used 1,545 composite T1 and T2 MR images (combined) in the axial view. The combined images were of 515 patients, each having the lower three levels (L3 to L5) of the lumbar spine. Broadly, our method is Generative Adversarial Network that we modified for our task. Our accuracy metric was the DICE Score. The accuracy was between 88.66 and 90.85%. We claim that our work is significant, as it establishes the ability of GAN to perform in-painting based on a slice for a disc, which indicates possibilities of future research on in-painting of larger stack of MR images and expansive synthetic dataset generation.
Subhajit Chakrabarty, Devesh Sarda
BIBM1
2024 A New Metric for Measuring Locational Health Access for Cancer Treatment
abstract
Ensuring access to cancer treatment facilities is essential for delivering timely care, yet various barriers such as geographic distance, socioeconomic factors, and social disparities can impede access in rural and urban regions. This study measured locational health access for colorectal cancer in the context of hospitals and population distribution in Louisiana. It used data of census tracts, hospital beds and providers, from the National Cancer Institute. By mapping the distribution of these healthcare facilities, the study revealed the potential of identifying significant challenges in accessing specialized cancer care. There is no existing locational health access metric in this domain. The contribution of this paper is that it meticulously calculated the actual road distance of each census tract centroid and each cancer-treating hospital, and offers a new locational health access metric. This metric considers the number of beds and number of oncologists, as a proxy for measurement of cancer treatment facilities. The significance of this work is that it can be applied in a larger scope (such as the country), with more variables, and for other diseases treated by hospitals. It has public policy implications; hospitals can be located through such data-driven analysis.
Subhajit Chakrabarty, Sweta Singh, Ismael Maya, Udaysinh Rathod, Debarshi Roy
BIBM1
2024 Geospatial Analysis of Socioeconomic Equity and Environmental Factors Influencing Lung Cancer Prevalence in US
abstract
This study investigates the complex interplay of socioeconomic and environmental factors contributing to lung cancer rates across US counties and parishes. The primary objectives are to predict lung cancer rates based on selected socioeconomic and environmental factors and to identify spatial clusters that reflect variations in these rates. Prior studies have explored how socioeconomic variables-such as income, insurance coverage, and smoking rates-along with environmental pollution (PM 2.5), influence lung cancer rates across different regions. This study employs a multi-stage methodology using geospatial data from various US parishes. We apply Multiscale Geographically Weighted Regression (MGWR), Spatial Autoregressive Models (SAR), and Kriging techniques to predict lung cancer rates, followed by K-means clustering based on predicted values. The results are visualized using Folium, and variable significance is assessed through P-values. Our analysis reveals several key findings. Smoking levels show a strong positive correlation with lung cancer rates, particularly in regions with higher smoking rates. Exposure to PM 2.5 significantly contributes to lung cancer risk, especially in areas with higher pollution levels. Poverty has a mixed effect, with some regions showing a negative association with lung cancer rates while others show minimal impact. Insurance coverage is linked with higher lung cancer rates in certain clusters, reflecting better access to healthcare or more comprehensive cancer reporting. Income does not show a strong or consistent effect across regions. This research provides crucial insights into lung cancer's socioeconomic equity and environmental determinants, suggesting targeted public health interventions to mitigate risks in vulnerable communities. By highlighting the significance of smoking and air quality, this study informs policymakers and health organizations on crafting effective strategies for lung cancer prevention tailored to specific regional contexts. The methodologies and findings from this study can be applied to other health-related issues, enabling further exploration of the socioeconomic equity impacts on various diseases and facilitating improved public health planning.
Mridula Mavuri, Subhajit Chakrabarty
BIBM2
2022 Automatic Image Segmentation of Monocytes and Index Computation Using Deep Learning
abstract
The classification of white cells plays an important part in medical diagnosis. The counts may suggest the presence of infection, inflammation, anemia, bleeding, and other blood-associated issues. More specifically, the counting in our study is the calculation of the Monocyte Index (MI). The purpose of MI is to determine whether the patient can receive units of blood by analyzing the assay. In case of incompatible blood transfusion, monocytes may ingest or adhere red cells. The index is the percentage of red cells adhered, ingested, or both, versus free monocytes. Manual methods for blood cell counting may take several hours and are highly prone to different sources of errors. Automatic methods, such as Linear Discriminant Analysis, Quadratic Discriminant Analysis, K-Nearest Neighbors, Naïve Bayes, Support Vector Machine, Convolutional Neural Network (CNN), Fast Region-based CNN, Faster Region-based CNN, Spatial Pyramidal Pooling network, Single Shot Detector and Mask Region-based CNN, exist for classification. However, these methods currently do not perform automatic counting and calculation of MI. The dataset is our own collection of images using ZEISS Axiocam 208 color/202 mono microscope camera. For the labels in our own collection, we performed polygonal annotation using the VGG Annotator tool. We trained the Mask R-CNN deep neural network model for automatic segmentation at the pixel-level, using COCO pre-trained weights. Our results look promising, as the Mask R-CNN can perform automatic segmentation with 72% accuracy. Compared to a medical laboratory scientist, the model can process large amount of data simultaneously, quickly and efficiently, with approximately the same judgment accuracy as a human eye. This may significantly reduce the burden of the laboratory scientist and provide a useful reference for doctors to identify a potential blood candidate to be transfused.
Luis A. Pena Marquez, Subhajit Chakrabarty
BIBM2
2019 A New Index for Measuring Inconsistencies in Independent Component Analysis Using Multi-sensor Data
Subhajit Chakrabarty, Haim Levkowitz
CDVE1
2019 Denoising and Stability using Independent Component Analysis in High Dimensions - Visual Inspection Still Required
abstract
Independent Component Analysis (ICA) has emerged as a useful method for separation of components, such as in removing noise from data. We examine one of the challenges of ICA - instability, particularly in high dimensions, when the independent components vary, each time when ICA is performed. This may be due to various causes including the stochastic nature of the algorithm and the additive noise. The objective of this study is to examine denoising and stability issues of ICA in high dimensions and make a comparative evaluation of select approaches. We take a challenging electrocardiogram dataset which is a high-dimensional time series of multiple sensors. We experiment with a mix of approaches and methods - for resampling, clustering, ICA algorithms and dimensionality. We check the internal validity using the Icasso stability index, the Amari separation performance index and the Minimum Distance (MD) index. The first key contribution of this work is that it finds counter-evidence to the claim that resampling (bootstrapping) tackles the question of stability. The second contribution is that it finds evidence of an important limitation of the Minimum Distance index when dealing with high dimensional data - the index may become highly concentrated and may remain sub-optimal at all dimensions. Selectively removing noise components by visual inspection can improve the Amari, the Icasso or the MD index values. Some automated tools exist, but in high dimensions, visual inspection of the individual components is still required for effective denoising - data driven methods are not good enough.
Subhajit Chakrabarty, Haim Levkowitz
IV (1)1
2018 Role of Prior Experience on Student Performance in the Introductory Undergraduate CS Course: (Abstract Only)
abstract
Student success rates in introductory computer science courses at colleges and universities across worldwide are scandalously low - 30% to 50% of students fail a first-semester course. At our university, over the past ten semesters, 40.6% of our students failed our first-semester computer science course. A survey (204 computer science students) was administered at the beginning of Fall 2016 to measure two hypothesized constructs: one on prior engagement in activities (such as summer camp, jobs) related to computer science, and another on prior experience in computer science topics. The prior experience bank consisted of several yes/no questions asking about familiarity with specific topics found in a first-year computer science course (e.g. globals, arrays, conditionals). The survey data was matched with the final course grades. Results revealed that the prior experience variables could be construed as a construct, but this was not the case with the prior engagement variables. We discovered a statistically significant relation between prior experience and course grade, with more experience predicting higher grades. Except ethnicity, other variables such as gender and transfer status were not found to be significant. This study emphasizes the need to consider the prior knowledge of students in building introductory computer science curricula, such as creating multiple tracks with students self-selecting into higher or lower prior-experience cohorts.
Subhajit Chakrabarty, Fred G. Martin
SIGCSE1
2018 The Tablet Game: An Embedded Assessment for Measuring Students' Programming Skill in App Inventor (Abstract Only)
abstract
Assessing students' learning of concepts in programming is an essential part of teaching computer science. We developed the Tablet Game, an embedded assessment that measures students' skill in identifying programming structures used to create various behaviors in MIT App Inventor. The assessment was implemented as an app for Android devices. Students conducted an activity in the app, and identified which code-blocks would create those behaviors. Students' responses were transmitted to our custom data-collection server. In two five-day app development summer camps held with middle school students, students completed the same Tablet Game assessment on day 1 and day 5. Students also completed pre/post surveys which gathered ethnographic data and asked about interest levels in computer science and prior programming experience. Using data from 44 students with pre/post assessments matched to surveys, our results indicated that (1) students with high self-reported prior experience in App Inventor outperformed students with low prior experience on the Tablet Game pre-test, indicating that the assessment measures programming skill and (2) students with low prior experience achieved equivalent results as the high prior experience cohort in the post-test, indicating that the camp was successful in imparting programming skills. Both of these results are statistically significant. Further, (3) there were no statistically significant differences in gender composition of the two experience cohorts, indicating that the camp was equally accessible to girls and boys.
Fred G. Martin, Chike Abuah, Subhajit Chakrabarty, Mark Sherman 0002, Diane Schilder
SIGCSE3