VLDB 2026 Research / reviewers in the wild / expert
Indrani Bhattacharya
dblp:172/7716
· DBLP profile ↗
14ranked-venue papers
7as first author
5since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 4 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 2 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 3 · 2 first-authorDatabases, data management, data science and information retrieval · 2 · 1 first-authorArtificial intelligence and machine learning · 1Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Sparse-XM: Spine Pose Adjustment with RGB-D Bone Segmentation via Cross-Modality Label Transfer
William R. Warner, Indrani Bhattacharya, Linton T. Evans, Sohail K. Mirza, Keith D. Paulsen, Xiaoyao Fan |
MICCAI (9) | 2 |
| 2025 | ProstAtlasDiff: Prostate cancer detection on MRI using Diffusion Probabilistic Models guided by population spatial cancer atlases
Cynthia Xinran Li, Indrani Bhattacharya, Sulaiman Vesal, Pejman Ghanouni, Hassan Jahanandish, Richard E. Fan, Geoffrey A. Sonn, Mirabela Rusu |
Medical Image Anal. | 2 |
| 2022 | Selective identification and localization of indolent and aggressive prostate cancers via CorrSigNIA: an MRI-pathology correlation and deep learning frameworkabstractAutomated methods for detecting prostate cancer and distinguishing indolent from aggressive disease on Magnetic Resonance Imaging (MRI) could assist in early diagnosis and treatment planning. Existing automated methods of prostate cancer detection mostly rely on ground truth labels with limited accuracy, ignore disease pathology characteristics observed on resected tissue, and cannot selectively identify aggressive (Gleason Pattern≥4) and indolent (Gleason Pattern=3) cancers when they co-exist in mixed lesions. In this paper, we present a radiology-pathology fusion approach, CorrSigNIA, for the selective identification and localization of indolent and aggressive prostate cancer on MRI. CorrSigNIA uses registered MRI and whole-mount histopathology images from radical prostatectomy patients to derive accurate ground truth labels and learn correlated features between radiology and pathology images. These correlated features are then used in a convolutional neural network architecture to detect and localize normal tissue, indolent cancer, and aggressive cancer on prostate MRI. CorrSigNIA was trained and validated on a dataset of 98 men, including 74 men that underwent radical prostatectomy and 24 men with normal prostate MRI. CorrSigNIA was tested on three independent test sets including 55 men that underwent radical prostatectomy, 275 men that underwent targeted biopsies, and 15 men with normal prostate MRI. CorrSigNIA achieved an accuracy of 80% in distinguishing between men with and without cancer, a lesion-level ROC-AUC of 0.81±0.31 in detecting cancers in both radical prostatectomy and biopsy cohort patients, and lesion-levels ROC-AUCs of 0.82±0.31 and 0.86±0.26 in detecting clinically significant cancers in radical prostatectomy and biopsy cohort patients respectively. CorrSigNIA consistently outperformed other methods across different evaluation metrics and cohorts. In clinical settings, CorrSigNIA may be used in prostate cancer detection as well as in selective identification of indolent and aggressive components of prostate cancer, thereby improving prostate cancer care by helping guide targeted biopsies, reducing unnecessary biopsies, and selecting and planning treatment. Indrani Bhattacharya, Arun Seetharaman, Christian Kunder, Wei Shao 0008, Leo C. Chen, Simon J. C. Soerensen, Jeffrey B. Wang, Nikola C. Teslovich, Richard E. Fan, Pejman Ghanouni, James D. Brooks, Geoffrey A. Sonn, Mirabela Rusu |
Medical Image Anal. | 1 |
| 2022 | Domain generalization for prostate segmentation in transrectal ultrasound images: A multi-center study
Sulaiman Vesal, Iani J. M. B. Gayo, Indrani Bhattacharya, Shyam Natarajan, Leonard S. Marks, Dean C. Barratt, Richard E. Fan, Yipeng Hu, Geoffrey A. Sonn, Mirabela Rusu |
Medical Image Anal. | 3 |
| 2021 | Weakly Supervised Registration of Prostate MRI and Histopathology Images
Wei Shao 0008, Indrani Bhattacharya, Simon J. C. Soerensen, Christian Kunder, Jeffrey B. Wang, Richard E. Fan, Pejman Ghanouni, James D. Brooks, Geoffrey A. Sonn, Mirabela Rusu |
MICCAI (4) | 2 |
| 2020 | CorrSigNet: Learning CORRelated Prostate Cancer SIGnatures from Radiology and Pathology Images for Improved Computer Aided DiagnosisabstractMagnetic Resonance Imaging (MRI) is widely used for screening and staging prostate cancer. However, many prostate cancers have subtle features which are not easily identifiable on MRI, resulting in missed diagnoses and alarming variability in radiologist interpretation. Machine learning models have been developed in an effort to improve cancer identification, but current models localize cancer using MRI-derived features, while failing to consider the disease pathology characteristics observed on resected tissue. In this paper, we propose CorrSigNet, an automated two-step model that localizes prostate cancer on MRI by capturing the pathology features of cancer. First, the model learns MRI signatures of cancer that are correlated with corresponding histopathology features using Common Representation Learning. Second, the model uses the learned correlated MRI features to train a Convolutional Neural Network to localize prostate cancer. The histopathology images are used only in the first step to learn the correlated features. Once learned, these correlated features can be extracted from MRI of new patients (without histopathology or surgery) to localize cancer. We trained and validated our framework on a unique dataset of 75 patients with 806 slices who underwent MRI followed by prostatectomy surgery. We tested our method on an independent test set of 20 prostatectomy patients (139 slices, 24 cancerous lesions, 1.12M pixels) and achieved a per-pixel sensitivity of 0.81, specificity of 0.71, AUC of 0.86 and a per-lesion AUC of \(0.96 \pm 0.07\), outperforming the current state-of-the-art accuracy in predicting prostate cancer using MRI. Indrani Bhattacharya, Arun Seetharaman, Wei Shao 0008, Rewa Sood, Christian Kunder, Richard E. Fan, Simon J. C. Soerensen, Jeffrey B. Wang, Pejman Ghanouni, Nikola C. Teslovich, James D. Brooks, Geoffrey A. Sonn, Mirabela Rusu |
MICCAI (2) | 1 |
| 2020 | Multiparty Visual Co-Occurrences for Estimating Personality Traits in Group MeetingsabstractParticipants’ body language during interactions with others in a group meeting can reveal important information about their individual personalities, as well as their contribution to a team. Here, we focus on the automatic extraction of visual features from each person, including her/his facial activity, body movement, and hand position, and how these features co-occur among team members (e.g., howfre- quently a person moves her/his arms or makes eye contact when she/he is the focus of attention of the group). We correlate these features with user questionnaires to reveal relationships with the "Big Five" personality traits (Openness, Conscientiousness, Extraversion, Agreeableness, Neuroti- cism), as well as with team judgements about the leader and dominant contributor in a conversation. We demonstrate that our algorithms achieve state-of-the-art accuracy with an average of 80% for Big-Five personality trait prediction, potentially enabling integration into automatic group meeting understanding systems. Lingyu Zhang 0002, Indrani Bhattacharya, Mallory Morgan, Michael Foley, Christoph Riedl, Brooke Foucault Welles, Richard J. Radke |
WACV | 2 |
| 2019 | Improved Visual Focus of Attention Estimation and Prosodic Features for Analyzing Group InteractionsabstractCollaborative group tasks require efficient and productive verbal and non-verbal interactions among the participants. Studying such interaction patterns could help groups perform more efficiently, but the detection and measurement of human behavior is challenging since it is inherently multimodal and changes on a millisecond time frame. In this paper, we present a method to study groups performing a collaborative decision-making task using non-verbal behavioral cues. First, we present a novel algorithm to estimate the visual focus of attention (VFOA) of participants using frontal cameras. The algorithm can be used in various group settings, and performs with a state-of-the-art accuracy of 90%. Secondly, we present prosodic features for non-verbal speech analysis. These features are commonly used in speech/music classification tasks, but are rarely used in human group interaction analysis. We validate our algorithms on a multimodal dataset of 14 group meetings with 45 participants, and show that a combination of VFOA-based visual metrics and prosodic-feature-based metrics can predict emergent group leaders with 64% accuracy and dominant contributors with 86% accuracy. We also report our findings on the correlations between the non-verbal behavioral metrics with gender, emotional intelligence, and the Big 5 personality traits. Lingyu Zhang 0002, Mallory Morgan, Indrani Bhattacharya, Michael Foley, Jonas Braasch, Christoph Riedl, Brooke Foucault Welles, Richard J. Radke |
ICMI | 3 |
| 2019 | Multimodal Dialog for Browsing Large Visual Catalogs using Exploration-Exploitation Paradigm in a Joint Embedding SpaceabstractWe present a multimodal dialog (MMD) system to assist online customers in visually browsing through large catalogs. Visual browsing allows customers to explore products beyond exact search results. We focus on a slightly asymmetric version of a complete MMD system, in that our agent can understand both text and image queries, but responds only in images. We formulate our problem of "showing the k best images to a user'', based on the dialog context so far, as sampling from a Gaussian Mixture Model (GMM) in a high dimensional joint multimodal embedding space. The joint embedding space is learned by Common Representation Learning and embeds both the text and the image queries. Our system remembers the context of the dialog, and uses an exploration-exploitation paradigm to assist in visual browsing. We train and evaluate the system on an MMD dataset that we synthesize from large catalog data. Our experiments and preliminary human evaluation show that the system is capable of learning and displaying relevant products with an average cosine similarity of 0.85 to the ground truth results, and is capable of engaging human users. Indrani Bhattacharya, Arkabandhu Chowdhury, Vikas C. Raykar |
ICMR | 1 |
| 2019 | The unobtrusive group interaction (UGI) corpusabstractStudying group dynamics requires fine-grained spatial and temporal understanding of human behavior. Social psychologists studying human interaction patterns in face-to-face group meetings often find themselves struggling with huge volumes of data that require many hours of tedious manual coding. There are only a few publicly available multi-modal datasets of face-to-face group meetings that enable the development of automated methods to study verbal and non-verbal human behavior. In this paper, we present a new, publicly available multi-modal dataset for group dynamics study that differs from previous datasets in its use of ceiling-mounted, unobtrusive depth sensors. These can be used for fine-grained analysis of head and body pose and gestures, without any concerns about participants' privacy or inhibited behavior. The dataset is complemented by synchronized and time-stamped meeting transcripts that allow analysis of spoken content. The dataset comprises 22 group meetings in which participants perform a standard collaborative group task designed to measure leadership and productivity. Participants' post-task questionnaires, including demographic information, are also provided as part of the dataset. We show the utility of the dataset in analyzing perceived leadership, contribution, and performance, by presenting results of multi-modal analysis using our sensor-fusion algorithms designed to automatically understand audio-visual interactions. Indrani Bhattacharya, Michael Foley, Christine Ku, Tongtao Zhang, Cameron Mine, Manling Li, Heng Ji 0001, Christoph Riedl, Brooke Foucault Welles, Richard J. Radke |
MMSys | 1 |
| 2018 | Unobtrusive Analysis of Group Interactions without CamerasabstractGroup meetings are often inefficient, unorganized and poorly documented. Factors including "group-think," fear of speaking, unfocused discussion, and bias can affect the performance of a group meeting. In order to actively or passively facilitate group meetings, automatically analyzing group interaction patterns is critical. Existing research on group dynamics analysis still heavily depends on video cameras in the lines of sight of participants or wearable sensors, both of which could affect the natural behavior of participants. In this thesis, we present a smart meeting room that combines microphones and unobtrusive ceiling-mounted Time-of-Flight (ToF) sensors to understand group dynamics in team meetings. Since the ToF sensors are ceiling-mounted and out of the lines of sight of the participants, we posit that their presence would not disrupt the natural interaction patterns of individuals. We collect a new multi-modal dataset of group interactions where participants have to complete a task by reaching a group consensus, and then fill out a post-task questionnaire. We use this dataset for the development of our algorithms and analysis of group meetings. In this paper, we combine the ceiling-mounted ToF sensors and lapel microphones to: (1) estimate the seated body orientation of participants, (2) estimate the head pose and visual focus of attention (VFOA) of meeting participants, (3) estimate the arm pose and body posture of participants, and (4) analyze the multimodal data for passive understanding of group meetings, with a focus on perceived leadership and contribution. Indrani Bhattacharya |
ICMI | 1 |
| 2018 | A Multimodal-Sensor-Enabled Room for Unobtrusive Group Meeting AnalysisabstractGroup meetings can suffer from serious problems that undermine performance, including bias, "groupthink", fear of speaking, and unfocused discussion. To better understand these issues, propose interventions, and thus improve team performance, we need to study human dynamics in group meetings. However, this process currently heavily depends on manual coding and video cameras. Manual coding is tedious, inaccurate, and subjective, while active video cameras can affect the natural behavior of meeting participants. Here, we present a smart meeting room that combines microphones and unobtrusive ceiling-mounted Time-of-Flight (ToF) sensors to understand group dynamics in team meetings. We automatically process the multimodal sensor outputs with signal, image, and natural language processing algorithms to estimate participant head pose, visual focus of attention (VFOA), non-verbal speech patterns, and discussion content. We derive metrics from these automatic estimates and correlate them with user-reported rankings of emergent group leaders and major contributors to produce accurate predictors. We validate our algorithms and report results on a new dataset of lunar survival tasks of 36 individuals across 10 groups collected in the multimodal-sensor-enabled smart room. Indrani Bhattacharya, Michael Foley, Tongtao Zhang, Christine Ku, Cameron Mine, Heng Ji 0001, Christoph Riedl, Brooke Foucault Welles, Richard J. Radke |
ICMI | 1 |
| 2016 | Arrays of single pixel time-of-flight sensors for privacy preserving tracking and coarse pose estimationabstractWe present a method for real-time person tracking and coarse pose estimation in a smart room using a sparse array of single pixel time-of flight (ToF) sensors mounted in the ceiling of the room. The single pixel sensors are relatively inexpensive compared to commercial ToF cameras and are privacy preserving in that they only return the range to a small set of hit points. The tracking algorithm includes higher level logic about how people move and interact in a room and makes estimates about the locations of people even in the absence of direct measurements. A maximum likelihood classifier based on features extracted from the time series of ToF measurements is used for robust pose classification into sitting, standing and walking states. We use both computer simulation and real-world experiments to show that the algorithms are capable of robust person tracking and pose estimation even with a sensor spacing of 60 cm (i.e., 1 sensor per ceiling tile). Indrani Bhattacharya, Richard J. Radke |
WACV | 1 |
| 2015 | Patient classification based on expanded query using 5-gram collocation and binary treeabstractPatients in rural India express their discomfort using keyword as query due to their lack of knowledge about the intended domain. Therefore, there is no scope of automatic revision of the query using feedback mechanism, unlike the existing query expansion methods. The paper aims at developing a primary level disease diagnosis system for the patients of rural India by expanding the query using 5-gram collocation model. First a string of five co-occurred words with respect to each query are obtained by consulting several medical documents. We call the string of five terms as bag of symptom (BoS), representing a concept. For each query there is multiple BoSs from which we select ten only based on their rank, measured using Log-likelihood ratio. However, all the terms in the BoS may not represent the disease symptom but semantically related concept similar to the symptoms of the symptom vocabulary (SV). We propose a novel binary tree based approach to calculate the degree of similarity (DoS) between the terms in the BoS and the symptoms in the SV using topology of the tree and term frequency-inverse document frequency (tf-idf) of the symptoms. The SV with respect to each BoS is encoded with DoS value and framed as feature vectors, which are mostly sparse. To remove sparsity in the feature vectors we apply singular value decomposition (SVD) method. Finally, the patients are classified into four probable diseases using 10-fold cross validation technique where the SV consists optimum no. of symptoms for such diseases. We classify pregnant women separately into two probable diseases and each of the cases the system shows satisfactory performance. Jaya Sil, Indrani Bhattacharya |
DSAA | 2 |