Anurag Das

dblp:190/4289 · DBLP profile ↗
← Back
14ranked-venue papers
9as first author
13since 2021 · last 2025
0009-0007-9800-0699ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 9 · 5 first-author · 8 since 2021Artificial intelligence and machine learning · 7 · 4 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 3 first-author · 4 since 2021
YearPublicationVenuePosition
2025 Health-Driven Personalized Metabolic Models of Postprandial Glucose Responses to Mixed Meals
Anurag Das, Ghady Nasrallah, Sicong Huang 0002, Bobak Mortazavi, Ricardo Gutierrez-Osuna
BSN1
2024 Macronutrient Constraints and Priors Improve Carbohydrate Predictions from Continuous Glucose Monitors
abstract
We propose an approach to estimate the macronu-trients in a meal automatically by analyzing the meal's glucose response using off-the-shelf wearable sensors (continuous glucose monitors). We rely on the fact that the shape of the glucose response to a meal depends on all the macronutrients in the meal, not just its carbohydrates (carbs). However, protein, fat, and fiber tend to affect the glucose response in similar ways, so recovering their individual amounts is numerically ill-conditioned. To address this problem, our approach compresses macronutrients into a latent variable that captures their correlated effects on glucose. Then, we train a machine learning model to predict the latent variable from the glucose response of a meal. Finally, we recover the amount of the original macronutrients by incorporating prior knowledge of how they co-occur in conventional meals. Using experimental data from 45 participants, we show that predicting carbs indirectly (through the latent variable) reduces the prediction error when compared to predicting carbs directly, i.e., without considering the protein and fats in the meal.
Anurag Das, Edmund Do, Namino Glanz, Wendy Bevier, Rony Santiago, David Kerr, Bobak Mortazavi, Ricardo Gutierrez-Osuna
BSN1
2024 MTA-CLIP: Language-Guided Semantic Segmentation with Mask-Text Alignment
Anurag Das, Xinting Hu, Li Jiang 0009, Bernt Schiele
ECCV (54)1
2024 Improving 2D Feature Representations by 3D-Aware Fine-Tuning
Yuanwen Yue, Anurag Das, Francis Engelmann, Siyu Tang 0001, Jan Eric Lenssen
ECCV (2)2
2024 Improving Mispronunciation Detection Using Speech Reconstruction
abstract
Training related machine learning tasks simultaneously can lead to improved performance on both tasks. Text- to-speech (TTS) and mispronunciation detection and diagnosis (MDD) both operate using phonetic information and we wanted to examine whether a boost in MDD performance can be by two tasks. We propose a network that reconstructs speech from the phones produced by the MDD system and computes a speech reconstruction loss. We hypothesize that the phones produced by the MDD system will be closer to the ground truth if the reconstructed speech sounds closer to the original speech. To test this, we first extract wav2vec features from a pre-trained model and feed it to the MDD system along with the text input. The MDD system then predicts the target annotated phones and then synthesizes speech from the predicted phones. The system is therefore trained by computing both a speech reconstruction loss as well as an MDD loss. Comparing the proposed systems against an identical system but without speech reconstruction and another state-of-the-art baseline, we found that the proposed system achieves higher mispronunciation detection and diagnosis (MDD) scores. On a set of sentences unseen during training, the and speaker verification simultaneously can lead to improve proposed system achieves higher MDD scores, which suggests that reconstructing the speech signal from the predicted phones helps the system generalize to new test sentences. We also tested whether the system can generate accented speech when the input phones have mispronunciations. Results from our perceptual experiments show that speech generated from phones containing mispronunciations sounds more accented and less intelligible than phones without any mispronunciations, which suggests that the system can identify differences in phones and generate the desired speech signal.
Anurag Das, Ricardo Gutierrez-Osuna
IEEE ACM Trans. Audio Speech Lang. Process.1
2023 Modeling the effect of non-exercise activity on peak post-prandial glucose in diabetes
abstract
The timing, intensity, and duration of postprandial exercise are important factors that reduce glucose excursions. When exercise is of moderate intensity, performed between 25 and 55 minutes after a meal, it results in greater attenuation of glucose. However, the potential glucose reduction for shorter-duration, non-exercise activity thermogenesis (NEAT) (such as activities of daily living) may also be beneficial, particularly in cases where exercise is neither feasible nor prudent. Therefore, we designed a system to capture blood glucose and activity intensity through internet of medical things devices and modeled the impact of the timing and duration of NEAT on peak glucose. This work designed a linear mixed effects model to evaluate the impact of NEAT on peak, postprandial glucose in a study of data captured on varied participants with or without diabetes. We found at least 25 minutes of NEAT starting 30 minutes after the meal most effectively reduced peak post-prandial glucose.Clinical Relevance— This work establishes the impact of NEAT on reducing post-prandial peak glucose in free-living environments as another method of controlling glucose surges
Edmund Do, Anurag Das, Namino Glanz, Wendy Bevier, Rony Santiago, David Kerr, Ricardo Gutierrez-Osuna, Bobak Mortazavi
BSN2
2023 Weakly-Supervised Domain Adaptive Semantic Segmentation with Prototypical Contrastive Learning
abstract
There has been a lot of effort in improving the performance of unsupervised domain adaptation for semantic segmentation task, however, there is still a huge gap in performance when compared with supervised learning. In this work, we propose a common framework to use different weak labels, e.g., image, point and coarse labels from the target domain to reduce this performance gap. Specifically, we propose to learn better prototypes that are representative class features by exploiting these weak labels. We use these improved prototypes for the contrastive alignment of class features. In particular, we perform two different feature alignments: first, we align pixel features with proto-types within each domain and second, we align pixel features from the source to prototype of target domain in an asymmetric way. This asymmetric alignment is beneficial as it preserves the target features during training, which is essential when weak labels are available from the target domain. Our experiments on various benchmarks show that our framework achieves significant improvement compared to existing works and can reduce the performance gap with supervised learning. Code will be available at https://github.com/anurag-198/WDASS.
Anurag Das, Yongqin Xian, Dengxin Dai, Bernt Schiele
CVPR1
2023 Decoupling Segmental and Prosodic Cues of Non-native Speech through Vector Quantization
Waris Quamer, Anurag Das, Ricardo Gutierrez-Osuna
INTERSPEECH2
2023 Urban Scene Semantic Segmentation with Low-Cost Coarse Annotation
abstract
For best performance, today’s semantic segmentation methods use large and carefully labeled datasets, requiring expensive annotation budgets. In this work, we show that coarse annotation is a low-cost but highly effective alternative for training semantic segmentation models. Considering the urban scene segmentation scenario, we lever-age cheap coarse annotations for real-world captured data, as well as synthetic data to train our model and show competitive performance compared with finely annotated real-world data. Specifically, we propose a coarse-to-fine self-training framework that generates pseudo labels for unlabeled regions of the coarsely annotated data, using synthetic data to improve predictions around the boundaries between semantic classes, and using cross-domain data augmentation to increase diversity. Our extensive experimental results on Cityscapes and BDD100k datasets demonstrate that our method achieves a significantly better performance vs annotation cost tradeoff, yielding a comparable performance to fully annotated data with only a small fraction of the annotation budget. Also, when used as pre-training, our framework performs better compared to the standard fully supervised setting.
Anurag Das, Yongqin Xian, Yang He 0005, Zeynep Akata, Bernt Schiele
WACV1
2022 Zero-Shot Foreign Accent Conversion without a Native Reference
Waris Quamer, Anurag Das, John Levis, Evgeny Chukharev-Hudilainen, Ricardo Gutierrez-Osuna
INTERSPEECH2
2022 Predicting the Macronutrient Composition of Mixed Meals From Dietary Biomarkers in Blood
abstract
Diet monitoring is an essential intervention component for a number of diseases, from type 2 diabetes to cardiovascular diseases. However, current methods for diet monitoring are burdensome and often inaccurate. In prior work, we showed that continuous glucose monitors (CGMs) may be used to predict meal macronutrients (e.g., carbohydrates, protein, fat) by analyzing the shape of the post-prandial glucose response. In this study, we examine a number of additional dietary biomarkers in blood by their ability to improve macronutrient prediction, compared to using CGMs alone. For this purpose, we conducted a nutritional study where (n = 10) participants consumed nine different mixed meals with varied but known macronutrient amounts, and we analyzed the concentration of 33 dietary biomarkers (including amino acids, insulin, triglycerides, and glucose) at various times post-prandially. Then, we built machine learning models to predict macronutrient amounts from (1) individual biomarkers and (2) their combinations. We find that the additional blood biomarkers provide complementary information, and more importantly, achieve lower normalized root mean squared error (NRMSE) for the three macronutrients (carbohydrates: 22.9%; protein: 23.4%; fat: 32.3%) than CGMs alone (carbohydrates: 28.9%, t(18) =1.64, p =0.060; protein: 46.4%, t(18) =5.38, p 0.001; fat: 40.0%, t(18) =2.09, p =0.025). Our main conclusion is that augmenting CGMs to measure these additional dietary biomarkers improves macronutrient prediction performance, and may ultimately lead to the development of automated methods to monitor nutritional intake. This work is significant to biomedical research as it provides a potential solution to the long-standing problem of diet monitoring, facilitating new interventions for a number of diseases.
Anurag Das, Bobak Mortazavi, Seyedhooman Sajjadi, Theodora Chaspari, Laura Ruebush, Nicolaas E. P. Deutz, Gerard L. Coté, Ricardo Gutierrez-Osuna
IEEE J. Biomed. Health Informatics1
2021 A Sparse Coding Approach to Automatic Diet Monitoring with Continuous Glucose Monitors
abstract
Measuring dietary intake is a major challenge in the management of chronic diseases. Current methods rely on self-report measures, which are cumbersome to obtain and often unreliable. This article presents an approach to estimate dietary intake automatically by analyzing the post-prandial glucose response (PPGR) of a meal, as measured with continuous glucose monitors. In particular, we propose a sparse-coding technique that can be used to estimate the amounts of macronutrients (carbohydrates, protein, fat) in a meal from the meal’s PPGR. We use Lasso regularization to represent the PPGR of a new meal as a sparse combination of PPGRs in a dictionary, then combine the sparse weights with the macronutrient amounts in the dictionary’s meals to estimate the macronutrients in the new meal. We evaluate the approach on a dataset containing nine standardized meals and their corresponding PPGRs, consumed by fifteen participants. The proposed technique consistently outperforms two baseline systems based on ridge regression and nearest-neighbors, in terms of correlation and normalized root mean square error of the predictions.
Anurag Das, Seyedhooman Sajjadi, Bobak Mortazavi, Theodora Chaspari, Projna Paromita, Laura Ruebush, Nicolaas E. P. Deutz, Ricardo Gutierrez-Osuna
ICASSP1
2021 Towards The Development of Subject-Independent Inverse Metabolic Models
abstract
Diet monitoring is an important component of interventions in type 2 diabetes, but is time intensive and often inaccurate. To address this issue, we describe an approach to monitor diet automatically, by analyzing fluctuations in glucose after a meal is consumed. In particular, we evaluate three standardization techniques (baseline correction, feature normalization, and model personalization) that can be used to compensate for the large individual differences that exist in food metabolism. Then, we build machine learning models to predict the amounts of macronutrients in a meal from the associated glucose responses. We evaluate the approach on a dataset containing glucose responses for 15 participants who consumed 9 meals. Three techniques improve the accuracy of the models: subtracting the baseline glucose, performing z-score normalization, and scaling the amount of macronutrients by each individuals’ body mass index.
Seyedhooman Sajjadi, Anurag Das, Ricardo Gutierrez-Osuna, Theodora Chaspari, Projna Paromita, Laura Ruebush, Nicolaas E. P. Deutz, Bobak Mortazavi
ICASSP2
2020 Understanding the Effect of Voice Quality and Accent on Talker Similarity
abstract
This paper presents a methodology to study the role of nonnative accents on talker recognition by humans. The methodology combines a state-of-the-art accent-conversion system to resynthesize the voice of a speaker with a different accent of her/his own, and a protocol for perceptual listening tests to measure the relative contribution of accent and voice quality on speaker similarity. Using a corpus of non-native and native speakers, we generated accent conversions in two different directions: non-native speakers with native accents, and native speakers with non-native accents. Then, we asked listeners to rate the similarity between 50 pairs of real or synthesized speakers. Using a linear mixed effects model, we find that (for our corpus) the effect of voice quality is five times as large as that of non-native accent, and that the effect goes away when speakers share the same (native) accent. We discuss the potential significance of this work in earwitness identification and sociophonetics.
Anurag Das, Guanlong Zhao, John Levis, Evgeny Chukharev-Hudilainen, Ricardo Gutierrez-Osuna
INTERSPEECH1