Rajlakshmi Guha

dblp:173/5333 · DBLP profile ↗
← Back
19ranked-venue papers
0as first author
13since 2021 · last 2025
0000-0002-4791-5182ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 7 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 4 since 2021Human-computer interaction and ubiquitous computing · 4 · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2025 INN-PAR: Invertible Neural Network for PPG to ABP Reconstruction
abstract
Non-invasive and continuous blood pressure (BP) monitoring is essential for the early prevention of many cardiovascular diseases. Estimating arterial blood pressure (ABP) from photoplethysmography (PPG) has emerged as a promising solution. However, existing deep learning approaches for PPG-to-ABP reconstruction (PAR) encounter certain information loss, impacting the precision of the reconstructed signal. To overcome this limitation, we introduce an invertible neural network for PPG to ABP reconstruction (INN-PAR), which employs a series of invertible blocks to jointly learn the mapping between PPG and its gradient with the ABP signal and its gradient. INN-PAR efficiently captures both forward and inverse mappings simultaneously, thereby preventing information loss. By integrating signal gradients into the learning process, INN-PAR enhances the network’s ability to capture essential high-frequency details, leading to more accurate signal reconstruction. Moreover, we propose a multi-scale convolution module (MSCM) within the invertible block, enabling the model to learn features across multiple scales effectively. We have experimented on two benchmark datasets, which show that INN-PAR significantly outperforms the state-of-the-art methods in both waveform reconstruction and BP measurement accuracy. Codes can be found at: https://github.com/soumitra1992/INNPAR-PPG2ABP.
Soumitra Kundu, Gargi Panda, Saumik Bhattacharya, Aurobinda Routray, Rajlakshmi Guha
ICASSP5
2025 A Study on The Impact of Foundation Models on Automatic Depression Detection from Speech Signals
abstract
An automatic depression detection (ADD) system using spoken language offers the opportunity to develop practical, low-cost tools to detect symptoms early. However, limited data availability, privacy concerns, and transcription efforts pose significant challenges. Recent advancements in foundational models, capable of understanding and processing multimodal inputs, present opportunities for enhancing ADD systems. This study explores various speech foundation models to investigate their impact on ADD. We leverage Whisper and MMS for automatic transcription and integrate speech and text embeddings into a language model optimized with low-rank adaptation (LoRA). In addition, we examine the effects of fine-tuning strategies and prompt formats on model performance. We used English and Bengali datasets to demonstrate the potential of our method in ADD, even with moderate-quality transcriptions. The best speech and language foundation models outperform baseline models on both datasets.
Bubai Maji, Monorama Swain, Shazia Nasreen, Debabrata Majumdar, Rajlakshmi Guha, Aurobinda Routray, Anders Søgaard
INTERSPEECH5
2024 The Eyes Have It: Exploring the Connection Between Domain-Knowledge and Perception
abstract
Leveraging prior domain-knowledge significantly impacts perception, and this study explores whether the prior domain-knowledge has any relation with the eye movements. We explore oculometrics of 95 graduate students from Architecture, Social Science, and Biology backgrounds, categorized by their domain-knowledge levels. Students were given unbiased (neutral), domain-biased (architecture and biology) tasks. Oculometrics of the students were captured during the tasks and were analyzed to determine if prior domain-knowledge affected eye movements. Results showed no significant difference between the three groups in the unbiased tasks. However, significant differences emerged in biased tasks between groups with and without domain-knowledge. Signature markers, such as peak saccadic velocity and mean pupil diameter, along with time-correlated parameters of total scanning duration, total fixation duration, total saccadic duration, and fixation count, were found to differentiate novices from those with prior domain-knowledge. The study highlights the importance of oculometrics to distinguish between novices and those with prior domain-knowledge.
Sonali Aatrai, Sandhya Gayatri Prabhala, Rajlakshmi Guha
ETRA4
2024 Learning through the Eyes: Unveiling Oculometrics for Domain-knowledge Identification
abstract
The process of learning, entailing the acquisition of new information, skills, and comprehension through study, experience, or teaching, is intricately linked to visual perception. Fusing visual experiences with cognitive processes plays a pivotal role in shaping and enriching domain-specific knowledge. In our research, we conducted experiments involving 102 participants categorized into specific domains, including 33 students from the Architecture group, 36 from the Mechanical group, and 33 from diverse domains (referred to as the Control group). Through domain-specific knowledge tasks in architecture and mechanical engineering, participants were classified based on their levels of expertise. Our findings reveal that significant eye markers are instrumental in discerning individuals with varying domain-specific knowledge. Notably, metrics such as Total Time Duration, Total Dwell Time, Number of Fixations, and Average Fixation Duration exhibit significance in distinguishing individuals across different domains, each manifesting at distinct time intervals. This understanding of the relevance of eye markers in identifying domain-specific knowledge has the potential to pave the way for innovative approaches in educational and training contexts. It may facilitate the development of personalized learning experiences tailored to individuals’ cognitive processes within specific domains, thereby optimizing knowledge acquisition and skill development.
Sonali Aatrai, Sandhya Gayatri Prabhala, Rajlakshmi Guha
ICALT3
2024 Hierarchical Classification of Frontotemporal Dementia Subtypes Utilizing Tabular-to-Image Data Conversion with Deep Learning Methods
Km Poonam, Venkata Sathwik Kotra, Rajlakshmi Guha, P. P. Chakrabarti 0001
ICPR (11)3
2024 Investigation of Layer-Wise Speech Representations in Self-Supervised Learning Models: A Cross-Lingual Study in Detecting Depression
Bubai Maji, Rajlakshmi Guha, Aurobinda Routray, Shazia Nasreen, Debabrata Majumdar
INTERSPEECH2
2024 Predicting Alzheimer's Disease Progression Using a Versatile Sequence-Length-Adaptive Encoder-Decoder LSTM Architecture
abstract
Detecting Alzheimer's disease (AD) accurately at an early stage is critical for planning and implementing disease-modifying treatments that can help prevent the progression to severe stages of the disease. In the existing literature, diagnostic test scores and clinical status have been provided for specific time points, and predicting the disease progression poses a significant challenge. However, few studies focus on longitudinal data to build deep-learning models for AD detection. These models are not stable to be relied upon in real medical settings due to a lack of adaptive training and testing. We aim to predict the individual's diagnostic status for the next six years in an adaptive manner where prediction performance improves with the number of patient visits. This study presents a Sequence-Length Adaptive Encoder-Decoder Long Short-Term Memory (SLA-ED LSTM) deep-learning model on longitudinal data obtained from the Alzheimer's Disease Neuroimaging Initiative archive. In the suggested approach, decoder LSTM dynamically adjusts to accommodate variations in training sequence length and inference length rather than being constrained to a fixed length. We evaluated the model performance for various sequence lengths and found that for inference length one, sequence length nine gives the highest average test accuracy and area under the receiver operating characteristic curves of 0.920 and 0.982, respectively. This insight suggests that data from nine visits effectively captures meaningful cognitive status changes and is adequate for accurate model training. We conducted a comparative analysis of the proposed model against state-of-the-art methods, revealing a significant improvement in disease progression prediction over the previous methods. Index Terms- Cognitive impairment, longitudinal data, multimodal data, encoder-decoder LSTM, progression predictionClinical relevanceThe proposed approach has the potential to improve understanding of Alzheimer's disease progression in diagnostics, facilitating early identification of various stages of cognitive decline leading to AD by considering its clinical variability.
Km Poonam, Rajlakshmi Guha, P. P. Chakrabarti 0001
IEEE J. Biomed. Health Informatics2
2023 Estimating Sub-categories of Cognitive Load: An Eye-tracking Study
Sonali Aatrai, Sparsh Kumar Jha, Rajlakshmi Guha
CogSci3
2023 Visual Perception and Performance: An Eye Tracking Study
abstract
This study explores the relationship between visual perception and performance. We investigate whether eye-metrics are consistent across various visual problem-solving tasks and if task complexity affects eye-metrics. Experiments were conducted on 102 participants using Tower of Hanoi (TOH), Image Sliding Puzzle (ISP) and 4 visual reasoning tasks with increasing complexity from the CLEVR dataset. Total Scanning Duration, Fixation count, Total Fixation Duration, and Total Saccadic Duration were found significant for distinguishing good and bad performers across tasks. Peak Velocity and Mean Pupil Diameter were found significant for varying task complexity. This was also reflected in time-matched samples of good and bad performers in TOH and ISP, though the content complexity of both these tasks remained constant. We propose that Peak Velocity and Mean Pupil Diameter are markers of ‘perceived task complexity’. Poor performers perceive tasks to be more complex even when content complexity is constant, and this affects their performance.
Sonali Aatrai, Sparsh Kumar Jha, Rajlakshmi Guha
ETRA3
2023 Automated Deep Learning Based Answer Generation to Psychometric Questionnaire: Mimicking Personality Traits
Anirban Lahiri, Shivam Raj, Utanko Mitra, Sunreeta Sen, Rajlakshmi Guha, Pabitra Mitra, P. P. Chakrabarti 0001, Anupam Basu
ICAART (3)5
2023 Multimodal Emotion Recognition Based on Deep Temporal Features Using Cross-Modal Transformer and Self-Attention
abstract
Multimodal speech emotion recognition (MSER) is an emerging and challenging field of research due to its more robust characteristics than unimodal. However, in multimodal approaches, the interactive relations for model building using different modalities of speech representations for emotion recognition have not been well investigated yet. To address this issue, we introduce a new approach to capturing the deep temporal features of audio and text. The audio features are learned with a convolution neural network (CNN) and a Bi-directional Gated Recurrent Unit (Bi-GRU) network. The textual features are represented by GloVe word embedding along with Bi-GRU. A cross-modal transformers block is designed for multimodal learning to capture better inter- and intra-interactions and temporal information between the audio and textual features. Further, a self-attention (SA) network is employed to select more important emotional information from the fused multimodal features. We evaluate the proposed method on the IEMOCAP dataset on four emotion classes (i.e., angry, neutral, sad, and happy). The proposed method performs significantly better than the most recent state-of-the-art MSER methods.
Bubai Maji, Monorama Swain, Rajlakshmi Guha, Aurobinda Routray
ICASSP3
2023 How Anxious Am 'Eye': An Eye Tracking Study
abstract
Anxiety is a psychological condition accompanied by various physiological changes along with cognitive ones. It can be manifested as a personality trait of being highly anxious or as a debilitating condition of Anxiety Neurosis. Literature shows individuals with high level of anxiety either attend or withdraw from threatening stimuli. This can be objectively tested by oculometric markers. In this study, healthy controls with high and moderate trait anxiety are compared with diagnosed cases of Anxiety Neurosis on their perception of ambiguous stimulus namely the Rorschach inkblot cards. Results suggest that the First Fixation Duration, Total Visit Count and Percentage fixated in a given area of interest are useful to distinguish clinical anxiety from a personality trait. Also, High Trait anxiety is significantly correlated with Anxiety neurosis in certain eye parameters. A trend is observed in oculometric changes with increasing anxiety.
Shazia Nasreen, Anup Kumar Roy, Rajlakshmi Guha, Debabrata Majumdar
IECON3
2022 A Study on Relative Performance of an Reinforcement Learning Agent and Human in a Psychometric Assessment Game
Utanko Mitra, Shivam Raj, Anirban Lahiri, Sunreeta Sen, Rajlakshmi Guha, Pabitra Mitra
CogSci5
2020 'Eye Can Reason'- How Eye Parameters Marked one's Performance in a Visual Reasoning Task
Kaustav Brahma, Pourush Sood, Rajlakshmi Guha, P. P. Chakrabarti 0001
CogSci3
2020 Antarjami: Exploring psychometric evaluation through a computer-based game
Anirban Lahiri, Utanko Mitra, Sunreeta Sen, Mreenal Chakraborty, Max Kleiman-Weiner, Rajlakshmi Guha, Pabitra Mitra, Anupam Basu, P. P. Chakrabarti 0001
CogSci6
2020 Improving the readability of dyslexic learners with mobile game-based sight-word training
abstract
Specific learning disabilities are a major obstacle in early learning processes and are a growing issue in India. Children with specific learning disabilities lack in necessary skills of reading and writing and require a personalized intervention by a clinical expert. However, the low expert-to-populace proportion is a noteworthy obstacle in effectively treating the disorder in the fully human-guided therapeutic set-up. The emerging use of Android handsets and effective e-learning technologies in the modern era provides ways to reduce the physical distance between the clinician and the child. This paper proposes one of the many solutions to help children with dyslexia in their early learning. This paper proposes an Android game-based intervention program to teach reading at the word level. The main objective of this work is to lessen the dependency of experts by providing a unified platform to assist both experts and children. The games are designed based on the sight words training a widely used and accepted intervention strategy. This work shows the developed prototype of the proposed approach, and the reviews from subject matter experts on the prototype through the Mobile Application Rating Scale.
Sajjad Ansari, Hirak Banerjee, Rajlakshmi Guha, Jayanta Mukhopadhyay
ICALT3
2020 A Novel Technique to Develop Cognitive Models for Ambiguous Image Identification Using Eye Tracker
abstract
Human behavior can be analyzed using Eye tracker. Thus, it is used for revealing the cognitive processes for object identification. Cognitive process is the mental ability for identification of what our eyes see. Vision with 20/20 sometimes may not reveal the purpose. In this study, ambiguous images are taken to observe the cognitive process in participants. During the perception of an object, a participant uses goal-directed search for identifying various objects. Dense gaze coordinates provide the region of interests and are considered as the target regions for object identification in ambiguous images. These data are used to develop cognitive models for identification of ambiguous images. Features such as, eye fixation, pupil diameter, fixation durations, moments of inertia, and polar moments are used for developing the cognitive model. Three different feature selection methods along with six different classifiers are used for the task of classification. The selection of a subset of features using hypothesis testing performed well, compared to principal component analysis based dimensionality reduction method. This study could be used in detecting whether a participants is lying or not while perceiving an ambiguous image.
Anup Kumar Roy, Md. Nadeem Akhtar, Manjunatha Mahadevappa, Rajlakshmi Guha, Jayanta Mukhopadhyay
IEEE Trans. Affect. Comput.4
2018 A Portable Personality Recognizer Based on Affective State Classification Using Spectral Fusion of Features
abstract
In this paper, we introduce a system named Portable Personality Recognizer (PPR), which classifies the personality of an individual using his/her transitions of affective states. This work attempts to reveal the latent relationship between emotions and personality of a person. Here, we train a hidden Markov model (HMM) with observable emotional states viz. Happiness (H), Anger (A), Surprise (S) and Disgust (D) and the hidden traits viz. Psychoticism (P), Extraversion (E) and Neuroticism (N). Based on the model, the system estimates the personality as Psychotic, Extravert or Neurotic. It does so by capturing the facial images of an individual using a visible and a thermal camera to decide the present affective state of the person. The emotion classification is carried out using fused eigenfeatures from the visible and blood perfused thermal images. The emotional state changes are observed using the trained HMM to estimate the personality. The proposed hardware prototype consists of a Banana Pi board with a seven inch LCD screen having a thermal and a visible camera add-ons. The system achieves an emotion classification accuracy of 87.145 percent, while an accuracy of 87.87 percent is achieved for personality recognition.
Anushree Basu, Anirban Dasgupta 0002, Anirud Thyagharajan, Aurobinda Routray, Rajlakshmi Guha, Pabitra Mitra
IEEE Trans. Affect. Comput.5
2017 The Indian Spontaneous Expression Database for Emotion Recognition
abstract
Automatic recognition of spontaneous facial expressions is a major challenge in the field of affective computing. Head rotation, face pose, illumination variation, occlusion etc. are the attributes that increase the complexity of recognition of spontaneous expressions in practical applications. Effective recognition of expressions depends significantly on the quality of the database used. Most well-known facial expression databases consist of posed expressions. However, currently there is a huge demand for spontaneous expression databases for the pragmatic implementation of the facial expression recognition algorithms. In this paper, we propose and establish a new facial expression database containing spontaneous expressions of both male and female participants of Indian origin. The database consists of 428 segmented video clips of the spontaneous facial expressions of 50 participants. In our experiment, emotions were induced among the participants by using emotional videos and simultaneously their self-ratings were collected for each experienced emotion. Facial expression clips were annotated carefully by four trained decoders, which were further validated by the nature of stimuli used and self-report of emotions. An extensive analysis was carried out on the database using several machine learning algorithms and the results are provided for future reference. Such a spontaneous database will help in the development and validation of algorithms for recognition of spontaneous expressions.
S. L. Happy, Priyadarshi Patnaik, Aurobinda Routray, Rajlakshmi Guha
IEEE Trans. Affect. Comput.4