Braja Gopal Patra

dblp:127/0064 · DBLP profile ↗
← Back
16ranked-venue papers
10as first author
6since 2021 · last 2025
0000-0003-2997-5314ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 8 · 4 first-author · 5 since 2021Artificial intelligence and machine learning · 4 · 4 first-authorDatabases, data management, data science and information retrieval · 4 · 2 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 since 2021
YearPublicationVenuePosition
2025 Extracting social support and social isolation information from clinical psychiatry notes: comparing a rule-based natural language processing system and a large language model
abstract
OBJECTIVES: Social support (SS) and social isolation (SI) are social determinants of health (SDOH) associated with psychiatric outcomes. In electronic health records (EHRs), individual-level SS/SI is typically documented in narrative clinical notes rather than as structured coded data. Natural language processing (NLP) algorithms can automate the otherwise labor-intensive process of extraction of such information. MATERIALS AND METHODS: Psychiatric encounter notes from Mount Sinai Health System (MSHS, n = 300) and Weill Cornell Medicine (WCM, n = 225) were annotated to create a gold-standard corpus. A rule-based system (RBS) involving lexicons and a large language model (LLM) using FLAN-T5-XL were developed to identify mentions of SS and SI and their subcategories (eg, social network, instrumental support, and loneliness). RESULTS: For extracting SS/SI, the RBS obtained higher macroaveraged F1-scores than the LLM at both MSHS (0.89 versus 0.65) and WCM (0.85 versus 0.82). For extracting the subcategories, the RBS also outperformed the LLM at both MSHS (0.90 versus 0.62) and WCM (0.82 versus 0.81). DISCUSSION AND CONCLUSION: Unexpectedly, the RBS outperformed the LLMs across all metrics. An intensive review demonstrates that this finding is due to the divergent approach taken by the RBS and LLM. The RBS was designed and refined to follow the same specific rules as the gold-standard annotations. Conversely, the LLM was more inclusive with categorization and conformed to common English-language understanding. Both approaches offer advantages, although additional replication studies are warranted.
Braja Gopal Patra, Lauren A. Lepow, Praneet Kasi Reddy Jagadeesh Kumar, Veer Vekaria, Mohit Manoj Sharma, Prakash Adekkanattu, Brian Fennessy, Gavin Hynes, Isotta Landi, Jorge A. Sanchez-Ruiz, Euijung Ryu, Joanna M. Biernacka, Girish N. Nadkarni, Ardesheer Talati, Myrna Weissman, Mark Olfson, J. John Mann, Yiye Zhang, Alexander Charney, Jyotishman Pathak
J. Am. Medical Informatics Assoc.1
2024 Identifying social determinants of health from clinical narratives: A study of performance, documentation ratio, and potential bias
Zehao Yu 0001, Cheng Peng 0009, Xi Yang 0015, Chong Dang, Prakash Adekkanattu, Braja Gopal Patra, Yifan Peng 0002, Jyotishman Pathak, Debbie L. Wilson, Ching-Yuan Chang, Wei-Hsuan Lo-Ciganic, Thomas J. George, William R. Hogan, Yi Guo 0005, Jiang Bian 0001, Yonghui Wu 0001
J. Biomed. Informatics6
2023 CoRL: A Cost-Responsive Learning Optimizer for Neural Networks
abstract
Selection of the optimal learning rate for training neural networks has often been a matter of concern for the machine learning community. The existing learning rates are dependent on multiple scaling factors. This paper proposes Cost-Responsive Learning (CoRL) which does not require manual hyper-parameter tuning. It maintains a linear relationship with the prediction error of the neural network. This is expected to offer the lowest learning rate at the global minima, and higher learning rates elsewhere. Hence, a number proportional to the prediction error is used as a learning rate, subject to the constraint that the number is within an acceptable range (here [0,1]). The derivation of an optimal learning rate from a given cost function is illustrated with the popular binary/categorical cross-entropy cost function(s). Experiments performed under multiple settings demonstrate that, with the CoRL optimizer needs no parameter tuning to obtain state-of-the-art results with significantly lower training time for equivalent performance.
Reshma Kar, Vijay Kumar Reddy Voddi, Braja Gopal Patra, Jyotishman Pathak
SMC3
2023 Scholarly recommendation systems: a literature survey
abstract
Abstract A scholarly recommendation system is an important tool for identifying prior and related resources such as literature, datasets, grants, and collaborators. A well-designed scholarly recommender significantly saves the time of researchers and can provide information that would not otherwise be considered. The usefulness of scholarly recommendations, especially literature recommendations, has been established by the widespread acceptance of web search engines such as CiteSeerX, Google Scholar, and Semantic Scholar. This article discusses different aspects and developments of scholarly recommendation systems. We searched the ACM Digital Library, DBLP, IEEE Explorer, and Scopus for publications in the domain of scholarly recommendations for literature, collaborators, reviewers, conferences and journals, datasets, and grant funding. In total, 225 publications were identified in these areas. We discuss methodologies used to develop scholarly recommender systems. Content-based filtering is the most commonly applied technique, whereas collaborative filtering is more popular among conference recommenders. The implementation of deep learning algorithms in scholarly recommendation systems is rare among the screened publications. We found fewer publications in the areas of the dataset and grant funding recommenders than in other areas. Furthermore, studies analyzing users’ feedback to improve scholarly recommendation systems are rare for recommenders. This survey provides background knowledge regarding existing research on scholarly recommenders and aids in developing future recommendation systems in this domain.
Braja Gopal Patra, Ashraf Yaseen, Rachit Sabharwal, Kirk Roberts, Tru Cao, Hulin Wu
Knowl. Inf. Syst.2
2021 Extracting Social Isolation Information From Psychiatric Notes in the Electronic Health Records
Lauren A. Lepow, Braja Gopal Patra, Isotta Landi, Prakash Adekkanattu, Jyotishman Pathak, Mark Olfson, J. John Mann, Euijung Ryu, Joanna M. Biernacka, Girish N. Nadkarni, Priya Wickramaratne, Myrna Weissman, Benjamin S. Glicksberg, Alexander Charney
AMIA2
2021 Extracting social determinants of health from electronic health records using natural language processing: a systematic review
abstract
OBJECTIVE: Social determinants of health (SDoH) are nonclinical dispositions that impact patient health risks and clinical outcomes. Leveraging SDoH in clinical decision-making can potentially improve diagnosis, treatment planning, and patient outcomes. Despite increased interest in capturing SDoH in electronic health records (EHRs), such information is typically locked in unstructured clinical notes. Natural language processing (NLP) is the key technology to extract SDoH information from clinical text and expand its utility in patient care and research. This article presents a systematic review of the state-of-the-art NLP approaches and tools that focus on identifying and extracting SDoH data from unstructured clinical text in EHRs. MATERIALS AND METHODS: A broad literature search was conducted in February 2021 using 3 scholarly databases (ACL Anthology, PubMed, and Scopus) following Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) guidelines. A total of 6402 publications were initially identified, and after applying the study inclusion criteria, 82 publications were selected for the final review. RESULTS: Smoking status (n = 27), substance use (n = 21), homelessness (n = 20), and alcohol use (n = 15) are the most frequently studied SDoH categories. Homelessness (n = 7) and other less-studied SDoH (eg, education, financial problems, social isolation and support, family problems) are mostly identified using rule-based approaches. In contrast, machine learning approaches are popular for identifying smoking status (n = 13), substance use (n = 9), and alcohol use (n = 9). CONCLUSION: NLP offers significant potential to extract SDoH data from narrative clinical notes, which in turn can aid in the development of screening tools, risk prediction models, and clinical decision support systems.
Braja Gopal Patra, Mohit Manoj Sharma, Veer Vekaria, Prakash Adekkanattu, Olga V. Patterson, Benjamin S. Glicksberg, Lauren A. Lepow, Euijung Ryu, Joanna M. Biernacka, Al'ona Furmanchuk, Thomas J. George, William R. Hogan, Yonghui Wu 0001, Xi Yang 0015, Jiang Bian 0001, Myrna Weissman, Priya Wickramaratne, J. John Mann, Mark Olfson, Thomas R. Campion Jr., Mark G. Weiner, Jyotishman Pathak
J. Am. Medical Informatics Assoc.1
2020 A content-based literature recommendation system for datasets to improve data reusability - A case study on Gene Expression Omnibus (GEO) datasets
Braja Gopal Patra, Vahed Maroufy, Babak Soltanalizadeh, W. Jim Zheng, Kirk Roberts, Hulin Wu
J. Biomed. Informatics1
2019 A Dataset Recommendation System for Researchers based on Publications
Braja Gopal Patra, Kirk Roberts, Hulin Wu
AMIA1
2018 Predicting Zika Prevention Techniques Discussed on Twitter: An Exploratory Study
abstract
Social media platforms are widely seen as a valuable medium to spread a wide range of information including charitable causes and health awareness. But given the flexibility provided by the social media platforms, it is important to ensure that the right kind of information is delivered to the right audience when needed. The pilot study presented in this paper considered a sample of Zika related tweets that were classified into different prevention techniques. The classification categories were drawn from the guidelines by CDC. Training a logistic regression model on the annotated data we found the accuracy to be 72%. The findings are significant in studying the effectiveness of social media platforms in spreading the right kind of information in time. This in turn can be useful in informing health care officials to take necessary steps with the help of real-time communication for such unfortunate events in future.
Soumik Mandal, Manasa Rath, Braja Gopal Patra
CHIIR4
2018 Multimodal mood classification of Hindi and Western songs
Braja Gopal Patra, Dipankar Das 0001, Sivaji Bandyopadhyay
J. Intell. Inf. Syst.1
2017 A Semantic Parsing Method for Mapping Clinical Questions to Logical Forms
Kirk Roberts, Braja Gopal Patra
AMIA2
2017 Labeling data and developing supervised framework for hindi music mood analysis
Braja Gopal Patra, Dipankar Das 0001, Sivaji Bandyopadhyay
J. Intell. Inf. Syst.1
2016 A Multilevel Approach to Sentiment Analysis of Figurative Language in Twitter
Braja Gopal Patra, Soumadeep Mazumdar, Dipankar Das 0001, Paolo Rosso, Sivaji Bandyopadhyay
CICLing (2)1
2016 Multimodal Mood Classification - A Case Study of Differences in Hindi and Western Songs
abstract
Music information retrieval has emerged as a mainstream research area in the past two decades. Experiments on music mood classification have been performed mainly on Western music based on audio, lyrics and a combination of both. Unfortunately, due to the scarcity of digitalized resources, Indian music fares poorly in music mood retrieval research. In this paper, we identified the mood taxonomy and prepared multimodal mood annotated datasets for Hindi and Western songs. We identified important audio and lyric features using correlation based feature selection technique. Finally, we developed mood classification systems using Support Vector Machines and Feed Forward Neural Networks based on the features collected from audio, lyrics, and a combination of both. The best performing multimodal systems achieved F-measures of 75.1 and 83.5 for classifying the moods of the Hindi and Western songs respectively using Feed Forward Neural Networks. A comparative analysis indicates that the selected features work well for mood classification of the Western songs and produces better results as compared to the mood classification systems for Hindi songs.
Braja Gopal Patra, Dipankar Das 0001, Sivaji Bandyopadhyay
COLING1
2015 Identifying Temporal Information and Tracking Sentiment in Cancer Patients' Interviews
Braja Gopal Patra, Nilabjya Ghosh, Dipankar Das 0001, Sivaji Bandyopadhyay
CICLing (2)1
2013 Construction of Emotional Lexicon Using Potts Model
Braja Gopal Patra, Hiroya Takamura, Dipankar Das 0001, Manabu Okumura, Sivaji Bandyopadhyay
IJCNLP1