Priya Rani

dblp:168/6689 · DBLP profile ↗
← Back
8ranked-venue papers
2as first author
8since 2021 · last 2025
0000-0002-2590-1185ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 2 first-author · 5 since 2021Software engineering, systems software and programming languages · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Safeguarding LLM-Applications: Specify or Train?
abstract
Large Language Models (LLMs) are powerful tools used in several applications such as conversational AI, and code generation. However, significant robustness concerns arise with LLMs in production, such as hallucinations, prompt injection attacks, harmful content generation, and challenges in maintaining accurate domain-specific content moderation. Guardrails aim to mitigate these challenges by aligning LLM outputs with desired behaviors without modifying the underlying models. Nvidia NeMo Guardrails, for instance, rely on specifying acceptable/unacceptable behaviours. However, it is challenging to predict and address potential issues of LLMs in advance to create these guardrails. Also, manual updates from software engineers are often required to maintain and refine these guardrails. We introduce LLM-Guards, specialised machine learning (ML) models trained to function as protective guards. Additionally, we present an automation pipeline for training and continual fine-tuning of these guards using reinforcement learning from human feedback (RLHF). We evaluated several small LLMs, including Llama-3, Mistral, and Gemma, as LLM-Guards for challenges such as moderation and detecting off-topic queries, and compared their performance against NeMo Guardrails. The proposed Llama-3 LLM-Guard outperformed NeMo Guardrails in detecting offtopic queries, achieving an accuracy of 98.7% compared to 81%. Furthermore, the LLM-Guard detected 97.86% of harmful queries” surpassing NeMo Guardrails by 19.86%.
Hala Abdelkader, Mohamed Almorsy, Sankhya Singh, Irini Logothetis, Priya Rani, Rajesh Vasa, Jean-Guy Schneider
CAIN5
2025 Prediction of Delirium Risk in Mild Cognitive Impairment Using Time-Series Data, Machine Learning and Comorbidity Patterns - A Retrospective Study
abstract
Delirium represents a significant clinical concern characterised by high morbidity and mortality rates, particularly in patients with mild cognitive impairment (MCI). This study investigates the associated risk factors for delirium by analysing the comorbidity patterns relevant to MCI and developing a longitudinal predictive model leveraging machine learning (ML) methodologies. A retrospective analysis utilising the MIMIC-IV v2.2 database was performed to evaluate comorbid conditions, survival probabilities, and predictive modelling outcomes. The examination of comorbidity patterns identified distinct risk profiles for the MCI population. Kaplan-Meier survival analysis demonstrated that individuals with MCI exhibit markedly reduced survival probabilities when developing delirium compared to their non-MCI counterparts, underscoring the heightened vulnerability within this cohort. For predictive modelling, a Long Short-Term Memory (LSTM) model was implemented utilising time-series data, demographic variables, Charlson Comorbidity Index (CCI) scores, and an array of comorbid conditions. The model demonstrated robust predictive capabilities with an Area Under the Receiver Operating Characteristic Curve (AUROC) of 0.92 and an Area Under the Precision-Recall Curve (AUPRC) of 0.91. This study underscores the critical role of comorbidities in evaluating delirium risk and highlights the efficacy of time-series predictive modeling in pinpointing patients at elevated risk for delirium development.
Santhakumar Ramamoorthy, Priya Rani, Glenn Matthews, Shaun L. Cloherty, Mahdi Babaei, James Mahon, Richard Kane, Christine Untario
IEEE J. Biomed. Health Informatics2
2024 ML-On-Rails: Safeguarding Machine Learning Models in Software Systems - A Case Study
abstract
Machine learning (ML), especially with the emergence of large language models (LLMs), has significantly transformed various industries. However, the transition from ML model prototyping to production use within software systems presents several challenges. These challenges primarily revolve around ensuring safety, security, and transparency, subsequently influencing the overall robustness and trustworthiness of ML models. In this paper, we introduce ML-On-Rails, a protocol designed to safeguard ML models, establish a well-defined endpoint interface for different ML tasks, and clear communication between ML providers and ML consumers (software engineers). ML-On-Rails enhances the robustness of ML models via incorporating detection capabilities to identify unique challenges specific to production ML. We evaluated the ML-On-Rails protocol through a real-world case study of the MoveReminder application. Through this evaluation, we emphasize the importance of safeguarding ML models in production.
Hala Abdelkader, Mohamed Almorsy, Scott Barnett, Jean-Guy Schneider, Priya Rani, Rajesh Vasa
CAIN5
2024 MaCmS: Magahi Code-mixed Dataset for Sentiment Analysis
abstract
The present paper introduces new sentiment data, MaCMS, for Magahi-Hindi-English (MHE) code-mixed language, where Magahi is a less-resourced minority language. This dataset is the first Magahi-Hindi-English code-mixed dataset for sentiment analysis tasks. Further, we also provide a linguistics analysis of the dataset to understand the structure of code-mixing and a statistical study to understand the language preferences of speakers with different polarities. With these analyses, we also train baseline models to evaluate the dataset’s quality.
Priya Rani, Theodorus Fransen, John P. McCrae, Gaurav Negi
LREC/COLING1
2024 Towards Robust ML-enabled Software Systems: Detecting Out-of-Distribution data using Gini Coefficients
abstract
Machine learning (ML) models have become essential components in software systems across several domains, such as autonomous driving, healthcare, and finance. The robustness of these ML models is crucial for maintaining the software systems performance and reliability. A significant challenge arises when these systems encounter out-of-distribution (OOD) data, examples that differ from the training data distribution. OOD data can cause a degradation of the software systems performance. Therefore, an effective OOD detection mechanism is essential for maintaining software system performance and robustness. Such a mechanism should identify and reject OOD inputs and alert software engineers. Current OOD detection methods rely on hyperparameters tuned with in-distribution and OOD data. However, defining the OOD data that the system will encounter in production is often infeasible. Further, the performance of these methods degrades with OOD data that has similar characteristics to the in-distribution data. In this paper, we propose a novel OOD detection method using the Gini coefficient. Our method does not require prior knowledge of OOD data or hyperparameter tuning. On common benchmark datasets, we show that our method outperforms the existing maximum softmax probability (MSP) baseline. For a model trained on the MNIST dataset, we improve the OOD detection rate by 4% on the CIFAR10 dataset and by more than 50% for the EMNIST dataset.
Hala Abdelkader, Jean-Guy Schneider, Mohamed Almorsy, Priya Rani, Rajesh Vasa
ASE4
2023 The Cardamom Workbench for Historical and Under-Resourced Languages
Adrian Doyle, Theodorus Fransen, Bernardo Stearns, John P. McCrae, Oksana Dereza, Priya Rani
LDK6
2022 Chatbots and Explainable Artificial Intelligence
abstract
For many areas of artificial intelligence, explainability provides assurance that a decision sits within an acceptable range of possible decisions. In the field of chatbots, however, the function of an AI is to provide an explanation to the user. Users may assume that the purpose of the chatbot is defined by this function. Before we consider the explanatory function of an AI chatbot, we should examine this assumption of purpose. In this research we consider two chatbot cases, the first being where the purpose may not be to inform the user, and secondly, where this should be the purpose. In the commercial sphere we identify two perspectives of AI chatbot purpose: that of the provider, and that of the user. No necessary commonality exists between these two perspectives of purpose. In the government services sphere, methods of increasing the alignment between requested information and appropriate response include “law as code” as a mechanism for simplifying the automation of regulation.
Marcus Wigan, Greg Adamson, Priya Rani, Nick Dyson, Fabian Horton
ISTAS3
2022 MHE: Code-Mixed Corpora for Similar Language Identification
abstract
This paper introduces a new Magahi-Hindi-English (MHE) code-mixed data-set for similar language identification (SMLID), where Magahi is a less-resourced minority language. This corpus provides a language id at two levels: word and sentence. This data-set is the first Magahi-Hindi-English code-mixed data-set for similar language identification task. Furthermore, we will discuss the complexity of the data-set and provide a few baselines for the language identification task.
Priya Rani, John P. McCrae, Theodorus Fransen
LREC1