VLDB 2026 Research / reviewers in the wild / expert
Kevin Desai
dblp:177/7859
· DBLP profile ↗
19ranked-venue papers
6as first author
12since 2021 · last 2026
0000-0002-2964-8981ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 17 · 6 first-author · 10 since 2021Human-computer interaction and ubiquitous computing · 5 · 5 since 2021Artificial intelligence and machine learning · 4 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SRA-Seg: Synthetic to Real Alignment for Semi-Supervised Medical Image Segmentation
OFM Riaz Rahman Aranya, Kevin Desai |
ICPR (1) | 2 |
| 2026 | TRACE: Temporal Radiology with Anatomical Change Explanation for Grounded X-Ray Report Generation
OFM Riaz Rahman Aranya, Kevin Desai |
ICPR (15) | 2 |
| 2025 | PatchFusionVR: Multitask Prediction of User Gaze, Reaction Time, and Cognitive Load in Virtual Reality from Multimodal SignalsabstractEnhancing user experience and performance, including task load in immersive environments, requires accurate prediction of user gaze point, reaction time, and mental and physical load uptake. Current gaze prediction approaches focus primarily on motion-based information, lacking physiological data, which leads to poor prediction accuracy in highly dynamic virtual reality (VR) environments. Traditional cognitive load measurements rely on post-task analysis without proper multimodal data integration and fail to capture the real-time dynamics of user states during interaction. Likewise, reaction time or attention load are often assessed only after the interaction, without using real-time immersive sensor data, which limits adaptive responsiveness. To tackle these limitations, we leveraged a comprehensive multimodal dataset - VRWalking, which recorded timestamped eye-tracking metrics, physiological signals (heart rate and galvanic skin response), and behavioral performance data during real-time engagement in a VR environment. We developed a unified multitask model based on the MultiPatchFormer architecture, which processes multimodal VR signals through dual patch projection branches for gaze and classification inputs. The model employs multiscale patch embeddings, cross-attention between gaze and classification pathways, channel attention, and transformer encoders to jointly predict continuous user gaze and classify reaction time, cognitive load (mental load and physical load). Our methodology achieved excellent predictive performance: 95.64% for reaction time, 98.01% for mental load, and 97.45% for physical load, with a MAPE (Mean Absolute Percentage Error) of 15.24% for gaze prediction. We applied Shapley Additive explanations (SHAP) analysis to interpret the model’s behavior across all features, including eye-tracking, head-tracking, and physiological signals. The analysis revealed which features most influenced the predictions of user gaze, reaction time, mental load, and physical load. Our methods, while based only on the VRWalking dataset, demonstrated strong performance across all tasks, suggesting promising potential for real-world VR applications such as interactive training systems that respond to user attention lapses, educational platforms that adapt to cognitive load, and performance assessments that consider physiological indicators. Md Irfan Pavel, M. Rasel Mahmud, Jyotirmay Nag Setu, Kevin Desai, John Quarles |
VRST | 4 |
| 2024 | EncodeNet: A Framework for Boosting DNN Accuracy with Entropy-Driven Generalized Converting Autoencoder
Hasanul Mahmud, Palden Lama, Kevin Desai, Sushil K. Prasad |
ICPR (2) | 3 |
| 2024 | Mazed and Confused: A Dataset of Cybersickness, Working Memory, Mental Load, Physical Load, and Attention During a Real Walking Task in VRabstractVirtual Reality (VR) is quickly establishing itself in various industries, including training, education, medicine, and entertainment, in which users are frequently required to carry out multiple complex cognitive and physical activities. However, the relationship between cognitive activities, physical activities, and familiar feelings of cybersickness is not well understood and thus can be unpredictable for developers. Researchers have previously provided labeled datasets for predicting cybersickness while users are stationary, but there have been few labeled datasets on cybersickness while users are physically walking. Moreover, it is unclear how walking while cybersick will affect cognitive load, even though room-scale interaction is typical in many VR games. Thus, from 39 participants, we collected head orientation, head position, eye tracking, images, physiological readings from external sensors, and the self-reported cybersickness severity, physical load, and mental load in VR. Throughout the data collection, participants navigated mazes via real walking and performed tasks challenging their attention and working memory. To demonstrate the dataset’s utility, we conducted a case study of training classifiers in which we achieved 95% accuracy for cybersickness severity classification. The noteworthy performance of the straightforward classifiers makes this dataset ideal for future researchers to develop cybersickness detection and reduction models. To better understand the features that helped with classification, we performed SHAP(SHapley Additive exPlanations) analysis, highlighting the importance of eye tracking and physiological measures for cybersickness prediction while walking. This open dataset can allow future researchers to study the connection between cybersickness and cognitive loads and develop prediction models. This dataset will empower future VR developers to design efficient and effective Virtual Environments by improving cognitive load management and minimizing cybersickness. Jyotirmay Nag Setu, Joshua M. Le, Ripan Kumar Kundu, Barry Giesbrecht, Tobias Höllerer, Khaza Anuarul Hoque, Kevin Desai, John Quarles |
ISMAR | 7 |
| 2024 | Investigating Personalization Techniques for Improved Cybersickness Prediction in Virtual Reality EnvironmentsabstractIn recent cybersickness research, there has been a growing interest in predicting cybersickness using real-time physiological data such as heart rate, galvanic skin response, eye tracking, postural sway, and electroencephalogram. However, the impact of individual factors such as age and gender, which are pivotal in determining cybersickness susceptibility, remains unknown in predictive models. Our research seeks to address this gap, underscoring the necessity for a more personalized approach to cybersickness prediction to ensure a better, more inclusive virtual reality experience. We hypothesize that a personalized cybersickness prediction model would outperform non-personalized models in predicting cybersickness. Evaluating this, we explored four personalization techniques: 1) data grouping, 2) transfer learning, 3) early shaping, and 4) sample weighing using an open-source cybersickness dataset. Our empirical results indicate that personalized models significantly improve prediction accuracy. For instance, with early shaping, the Deep Temporal Convolutional Neural Network (DeepTCN) model achieved a 69.7% reduction in RMSE compared to its non-personalized version. Our study provides evidence of personalization techniques' benefits in improving cybersickness prediction. These findings have implications for developing personalized cybersickness prediction models tailored to individual differences, which can be used to develop personalized cybersickness reduction techniques in the future. Umama Tasnim, Rifatul Islam, Kevin Desai, John Quarles |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2022 | Virtepex: Virtual Remote Tele-Physical Examination SystemabstractRemote strength assessment is critical for providing accessible rehabilitation, especially in the absence of in-person meetings due to the pandemic. In this paper, we introduce ”Virtepex”, an immersive exergame for remote strength assessment developed through participatory design principles. We bring out the design process starting with a needs assessment to highlight the challenges for physicians in telehealth, followed by the expert guidelines for iterative system refinement. Virtepex addresses the challenges for remote strength assessment through a marker-less and an easy-to-setup strength estimation pipeline. It utilizes an RGB-D camera for motion tracking and an inverse dynamics module for force estimation. The force estimates are used for VR object interaction and can be assessed by a physician synchronously or asynchronously for an objective evaluation. Validation by external experts shows that Virtepex produces reliable force estimates for upper body joints, indicating the potential of marker-less force estimation for future remote assessment designs. Ninad Khargonkar, Kevin Desai, B. Prabhakaran 0001, Thiru Annaswamy |
Conference on Designing Interactive Systems | 2 |
| 2022 | PMPNet: Pixel Movement Prediction Network for Monocular Depth Estimation in Dynamic ScenesabstractIn this paper, we propose a novel method for monocular depth estimation in dynamic scenes. We first explore the arbitrariness of object’s movement trajectory in dynamic scenes theoretically. To overcome the arbitrariness, we use assume that points move along a straight line over short distances and then summarize it as a triangular constraint loss in two dimensional Euclidean space. This triangular loss function is used as part of our proposed pixel movement prediction network, PMPNet, to estimate a dense depth map from a single input image. To overcome the depth inconsistency problem around the edges, we propose a deformable support window module that learns features from different shapes of objects, making depth value more accurate around edge area. The proposed model is trained and tested on two outdoor datasets - KITTI and Make3D, as well as an indoor dataset - NYU Depth V2. The quantitative and qualitative results reported on these datasets demonstrate the success of our proposed model when compared against other approaches. Ablation study results on the KITTI dataset also validate the effectiveness of the proposed pixel movement prediction module as well as the deformable support window module. Kebin Peng, John Quarles, Kevin Desai |
ICPR | 3 |
| 2022 | Towards Forecasting the Onset of Cybersickness by Fusing Physiological, Head-tracking and Eye-tracking with Multimodal Deep Fusion NetworkabstractA plethora of studies has been conducted to detect and reduce cybersickness in real-time. However, prior attempts to detect and minimize cybersickness after its onset may be ineffective as the onset tends to persist beyond its first occurrence. By forecasting the onset of cybersickness, it may be possible to mitigate the severity of cybersickness through earlier interventions. This research proposed a multimodal deep fusion approach to forecast cybersickness from the user’s physiological, head-tracking, and eye-tracking data. We proposed several hybrid multimodal deep fusion neural networks with Long short-term memory (LSTMs), Neural basis expansion analysis for interpretable time series forecasting(NBEATs) and Deep Temporal Convolutional Networks(DeepTCN) neural models to forecast cybersickness 30-60s in advance to its onset. To validate our proposed approach, we recruited 30 participants who were immersed in five virtual reality simulations. We collected eye-tracking, head-tracking, heart rate, and galvanic skin response data and used the fast-motion scale as ground truth. Our results suggest that the DeepTCN model with our proposed multimodal fusion network can forecast cybersickness onset 60 seconds in advance with a root-mean-square error of 0.49 (on a scale from 0-10). Furthermore, our results demonstrated that fusing eye tracking, heart rate, and galvanic skin response data outperformed other data fusion approaches. This research clarifies how early cybersickness can be forecast, paving the way for future research on early cybersickness mitigation approaches. Rifatul Islam, Kevin Desai, John Quarles |
ISMAR | 2 |
| 2021 | Generalized Zero-Shot Learning Using Multimodal Variational Auto-Encoder With Semantic ConceptsabstractWith the ever-increasing amount of data, the central challenge in multimodal learning involves limitations of labelled samples For the task of classification, techniques such as meta-learning, zero-shot learning, and few-shot learning showcase the ability to learn information about novel classes based on prior knowledge. Recent techniques try to learn a cross-modal mapping between the semantic space and the image space. However, they tend to ignore the local and global semantic knowledge. To overcome this problem, we propose a Multimodal Variational Auto-Encoder (M-VAE) which can learn the shared latent space of image features and the semantic space. In our approach we concatenate multimodal data to a single embedding before passing it to the VAE for learning the latent space. We propose the use of a multi-modal loss during the reconstruction of the feature embedding through the decoder. Our approach is capable to correlating modalities and exploit the local and global semantic knowledge for novel sample predictions. Our experimental results using a MLP classifier on four benchmark datasets show that our proposed model outperforms the current state-of-the-art approaches for generalized zero-shot learning. Nihar Bendre, Kevin Desai, Peyman Najafirad |
ICIP | 2 |
| 2021 | Cybersickness Prediction from Integrated HMD's Sensors: A Multimodal Deep Fusion Approach using Eye-tracking and Head-tracking DataabstractCybersickness prediction is one of the significant research challenges for real-time cybersickness reduction. Researchers have proposed different approaches for predicting cybersickness from bio-physiological data (e.g., heart rate, breathing rate, electroencephalogram). However, collecting bio-physiological data often requires external sensors, limiting locomotion and 3D-object manipulation during the virtual reality (VR) experience. Limited research has been done to predict cybersickness from the data readily available from the integrated sensors in head-mounted displays (HMDs) (e.g., head-tracking, eye-tracking, motion features), allowing free locomotion and 3D-object manipulation. This research proposes a novel deep fusion network to predict cybersickness severity from heterogeneous data readily available from the integrated HMD sensors. We extracted 1755 stereoscopic videos, eye-tracking, and head-tracking data along with the corresponding self-reported cybersickness severity collected from 30 participants during their VR gameplay. We applied several deep fusion approaches with the heterogeneous data collected from the participants. Our results suggest that cybersickness can be predicted with an accuracy of 87.77% and a root-mean-square error of 0.51 when using only eye-tracking and head-tracking data. We concluded that eye-tracking and head-tracking data are well suited for a standalone cybersickness prediction framework. Rifatul Islam, Kevin Desai, John Quarles |
ISMAR | 2 |
| 2021 | Show Why the Answer is Correct! Towards Explainable AI using Compositional Temporal AttentionabstractVisual Question Answering (VQA) models have achieved significant success in recent times. Despite the success of VQA models, they are mostly black-box models providing no reasoning about the predicted answer, thus raising questions for their applicability in safety-critical such as autonomous systems and cyber-security. Current state of the art fail to better complex questions and thus are unable to exploit compositionality. To minimize the black-box effect of these models and also to make them better exploit compositionality, we propose a Dynamic Neural Network (DMN), which can understand a particular question and then dynamically assemble various relatively shallow deep learning modules from a pool of modules to form a network. We incorporate compositional temporal attention to these deep learning based modules to increase compositionality exploitation. This results in achieving better understanding of complex questions and also provides reasoning as to why the module predicts a particular answer. Experimental analysis on the two benchmark datasets, VQA2.0 and CLEVR, depicts that our model outperforms the previous approaches for Visual Question Answering task as well as provides better reasoning, thus making it reliable for mission critical applications like safety and security. Nihar Bendre, Kevin Desai, Peyman Najafirad |
SMC | 2 |
| 2019 | Using Mr. MAPP for Lower Limb Phantom Pain ManagementabstractPhantom pain is a chronic pain that is experienced as a vivid sensation stemming from the missing limb. From traditional mirror box to virtual reality-based approaches, a wide spectrum of treatments using mimic feedback of the amputated limb have been developed for alleviating phantom limb pain. In our previous work, Mixed reality-based framework for MAnaging Phantom Pain (Mr.MAPP) was presented and used to generate a virtual phantom upper limb, in real time, to manage the phantom pain. However, amputation of the lower limb is more common than that of the upper limb. Hence, in this paper, on top of demonstrating the reproducibility of the Mr.MAPP framework for upper limb, we extend it to manage lower limb phantom pain as well. Unlike an upper limb amputee, a patient with lower limb amputated is constrained to perform the training procedure in a sitting posture. Accordingly, virtual training games are designed for lower limb exercises with sitting posture such as knee flexion and extension, ankle dorsiflexion and tandem coordinated movement. Finally, the technical details of the system setup for playing the training games are introduced. Kanchan Bahirat, Yu-Yen Chung, Thiru Annaswamy, Gargi Raval, Kevin Desai, B. Prabhakaran 0001, Michael Riegler 0001 |
ACM Multimedia | 5 |
| 2018 | Combining skeletal poses for 3D human model generation using multiple kinectsabstractRGB-D cameras, such as the Microsoft Kinect, provide us with the 3D information, color and depth, associated with the scene. Interactive 3D Tele-Immersion (i3DTI) systems use such RGB-D cameras to capture the person present in the scene in order to collaborate with other remote users and interact with the virtual objects present in the environment. Using a single camera, it becomes difficult to estimate an accurate skeletal pose and complete 3D model of the person, especially when the person is not in the complete view of the camera. With multiple cameras, even with partial views, it is possible to get a more accurate estimate of the skeleton of the person leading to a better and complete 3D model. In this paper, we present a real-time skeletal pose identification approach that leverages on the inaccurate skeletons of the individual Kinects, and provides a combined optimized skeleton. We estimate the Probability of an Accurate Joint (PAJ) for each joint from all of the Kinect skeletons. We determine the correct direction of the person and assign the correct joint sides for each skeleton. We then use a greedy consensus approach to combine the highly probable and accurate joints to estimate the combined skeleton. Using the individual skeletons, we segment the point clouds from all the cameras. We use the already computed PAJ values to obtain the Probability of an Accurate Bone (PAB). The individual point clouds are then combined one segment after another using the calculated PAB values. The generated combined point cloud is a complete and accurate 3D representation of the person present in the scene. We validate our estimated skeleton against two well-known methods by computing the error distance between the best view Kinect skeleton and the estimated skeleton. An exhaustive analysis is performed by using around 500000 skeletal frames in total, captured using 7 users and 7 cameras. Visual analysis is performed by checking whether the estimated skeleton is completely present within the human model. We also develop a 3D Holo-Bubble game to showcase the real-time performance of the combined skeleton and point cloud. Our results show that our method performs better than the state-of-the-art approaches that use multiple Kinects, in terms of objective error, visual quality and real-time user performance. Kevin Desai, B. Prabhakaran 0001, Suraj Raghuraman |
MMSys | 1 |
| 2018 | Skeleton-based continuous extrinsic calibration of multiple RGB-D kinect camerasabstractApplications involving 3D scanning and reconstruction & 3D Tele-immersion provide an immersive experience by capturing a scene using multiple RGB-D cameras, such as Kinect. Prior knowledge of intrinsic calibration of each of the cameras, and extrinsic calibration between cameras, is essential to reconstruct the captured data. The intrinsic calibration for a given camera rarely ever changes, so only needs to be estimated once. However, the extrinsic calibration between cameras can change, even with a small nudge to the camera. Calibration accuracy depends on sensor noise, features used, sampling method, etc., resulting in the need for iterative calibration to achieve good calibration. Kevin Desai, B. Prabhakaran 0001, Suraj Raghuraman |
MMSys | 1 |
| 2017 | Learning-based objective evaluation of 3D human open meshesabstractCurrent state-of-the-art mesh quality measures evaluate closed and complete meshes obtained after mesh postprocessing applications, such as mesh simplification or watermarking, and compare them against the corresponding reference mesh. Emerging 3D immersive VR/AR applications use noisy 3D point cloud, typically from single RGB-D camera (such as Microsoft's Kinect) to generate standalone (no reference) 3D human open mesh (with boundaries) in real time, that needs evaluation. A learning-based objective measure is proposed to rate the visual quality by emulating human perception of 3D human open mesh quality. 2-pronged objective evaluation is performed: (a) Global holistic score captures the efficacy of the mesh to represent the human model as a whole, by considering mesh completeness and mesh noise. (b) Local part-based score caters to the need of varying roughness in different parts of the human body, by finding the deviation in the face normals for all the adjacent triangles in that part (segment). Learning technique aligns the objective scores with the subjective user evaluation, in turn combining the concepts of white-box and black-box evaluation for 3D meshes. Experimental results for a database, specifically generated for the purpose proves the efficacy of the proposed method. Kevin Desai, Kanchan Bahirat, B. Prabhakaran 0001 |
ICME | 1 |
| 2017 | QoE Studies on Interactive 3D Tele-ImmersionabstractUsers' Quality of Experience (QoE) in Interactive 3D Tele-Immersion (i3DTI) systems is influenced by several factors such as the quality of the "live" 3D avatars of the users, network latency, rendering methodology (head mounted display or regular TV type of display), etc. Hence, it becomes important to answer the question: "Is Visual Quality (VQ) the only factor to be considered or do better immersion and faster interactions matter, in having good QoE?" To answer this question, in this paper, a highly optimized state-of-the-art i3DTI framework implementation is introduced along with a soccer-inspired penalty shootout game. This game allows users to experience various situations, view different angles, perceive delays, track virtual ball motion, and play naturally using their entire bodies. A head mounted display device - Oculus Rift allows the users to get completely immersed and perform better in the penalty shootout game compared to watching themselves play on a 3D TV. This scenario is obtained in a controlled lab setting with ultra-high-speed network that has ultra-low latency, high VQ, fast and realistic interactions. Such ideal conditions are not typically available in a wide area Internet. Hence, for faster interaction scenario, we used lesser RGB-D cameras for 3D reconstruction, thereby reducing the model quality significantly. The high responsiveness of the game masked the user's perception of quality, and resulted in them not noticing the lower VQ of the reconstructed 3D models. Based on the results from the user study that focused on immersion, interaction and VQ aspects; users felt the game to be visually appealing, intuitive, engaging, and highly entertaining in all of the scenarios. Kevin Desai, Suraj Raghuraman, Rong Jin 0003, B. Prabhakaran 0001 |
ISM | 1 |
| 2016 | Augmented reality-based exergames for rehabilitationabstractRehabilitation for stroke afflicted patients, through exercises tailored for individual needs, aims at relearning basic motor skills, especially in the extremities. Rehabilitation through Augmented Reality (AR) based games engage and motivate patients to perform exercises which, otherwise, maybe boring and monotonic. Also, mirror therapy allows users to observe one's own movements in the game providing them with good visual feedback. This paper presents an augmented reality based system for rehabilitation by playing four interactive, cognitive and fun Exergames (exercise and gaming). Kevin Desai, Kanchan Bahirat, Sudhir Ramalingam, B. Prabhakaran 0001, Thiru Annaswamy, Una E. Makris |
MMSys | 1 |
| 2015 | Network Adaptive Textured Mesh Generation for Collaborative 3D Tele-Immersionabstract3D Tele-Immersion (3DTI) has emerged as an efficient environment for virtual interactions and collaborations in a variety of fields like rehabilitation, education, gaming, etc. In 3DTI, geographically distributed users are captured using multiple cameras and immersed in a single virtual environment. The quality of experience depends on the available network bandwidth, quality of the 3D model generated and the time taken for rendering. In a collaborative environment, achieving high quality, high frame rate rendering by transmitting data to multiple sites having different bandwidth is challenging. In this paper we introduce a network adaptive textured mesh generation scheme to transmit varying quality data based on the available bandwidth. To reduce the volume of information transmitted, a visual quality based vertex selection approach is used to generate a sparse representation of the user. This sparse representation is then transmitted to the receiver side where a sweep-line based technique is used to generate a 3D mesh of the user. High visual quality is maintained by transmitting a high resolution texture image compressed using a lossy compression algorithm. In our studies users were unable to notice visual quality variations of the rendered 3D model even at 90% compression. Kevin Desai, Kanchan Bahirat, Suraj Raghuraman, B. Prabhakaran 0001 |
ISM | 1 |