Preethu Rose Anish

dblp:121/3938 · also Preethu Anish, Preethu Rose · DBLP profile ↗
← Back
20ranked-venue papers
6as first author
13since 2021 · last 2026
0009-0001-7279-8993ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 17 · 6 first-author · 10 since 2021Artificial intelligence and machine learning · 4 · 4 since 2021
YearPublicationVenuePosition
2026 Mitigating Misinterpretation in Policy Documents through Automated Language Understanding
Momojit Biswas, Anka Chandrahas Tummepalli, Preethu Rose Anish
LREC3
2026 C-PASS: An organization-centric framework for Compliance and Penalty Assessment using Large Language Models
Gokul Rejithkumar, Sachin Pawar, Pavithra P. M. Nair, Nitin Ramrakhiyani, Preethu Rose Anish
Inf. Softw. Technol.5
2025 Exploring Zero-Shot App Review Classification with ChatGPT: Challenges and Potential
abstract
App reviews are a critical source of user feedback, offering valuable insights into an app’s performance, features, usability, and overall user experience. Effectively analyzing these reviews is essential for guiding app development, prioritizing feature updates, and enhancing user satisfaction. Classifying reviews into functional and non-functional requirements play a pivotal role in distinguishing feedback related to specific app features (functional requirements) from feedback concerning broader quality attributes, such as performance, usability, and reliability (non-functional requirements). Both categories are integral to informed development decisions. Traditional approaches to classifying app reviews are hindered by the need for large, domain-specific datasets, which are often costly and time-consuming to curate. This study explores the potential of zero-shot learning with ChatGPT for classifying app reviews into four categories: functional requirement, non-functional requirement, both, or neither. We evaluate ChatGPT’s performance on a benchmark dataset of 1,880 manually annotated reviews from ten diverse apps spanning multiple domains. Our findings demonstrate that ChatGPT achieves a robust F1 score of 0.842 in review classification, despite certain challenges and limitations. Additionally, we examine how factors such as review readability and length impact classification accuracy and conduct a manual analysis to identify review categories more prone to misclassification.
Mohit Chaudhary, Preethu Rose Anish
EASE3
2025 Exploring LLMs for Stakeholder-Specific Insight Generation from Software Contracts
abstract
Background: Software contracts are legally binding agreements that define terms for development, licensing, usage, and distribution. Despite their importance for compliance and stakeholder clarity, their complexity often impedes readability, and generic clause summarization fails to capture stakeholder-specific responsibilities, especially when clauses span multiple departments. Aim: This study explores the use of large language models (LLMs) to generate stakeholder-specific insights from complex contractual clauses in software engineering. Method: We evaluate zero-shot, few-shot prompting, and finetuning techniques across open-source models including T5, Llama 3.1 & 3.2, PEGASUS, BART, Mistral, Gemma 3, and Qwen 2.5. Commercial models like ChatGPT and Claude are excluded due to data privacy constraints, focusing instead on deployable models suitable for proprietary datasets. Using a dataset of 4,000 contractual clauses, we assess performance through ROUGE, METEOR, BLEU scores, and human-evaluation metrics such as fluency, coherence, informativeness, and relevance. Results: Finetuning emerged as the most effective technique, improving performance by$150-200 {\%}$over others. Among all models, the finetuned LlaMa 3.2 achieved the highest scores-over 0.9 on quantitative metrics, consistent ‘High’ ratings across qualitative metrics, and sub-second insight generation. Conclusions: The finetuned Llama 3.2 model demonstrates strong practical applicability and has been successfully integrated into the Software Contracts Governance System (SCGS) of a major IT vendor, validating its utility in real-world contract analysis.
Jyoti S. Shukla, Aditya Kahol, Mohit Chaudhary, Preethu Rose Anish
ESEM4
2025 Optimized Domain-Specific Text Processing with Keyword Knowledge Distillation (KKD)
abstract
The generative and reasoning capabilities of Pre-Trained Language Models (PLMs) have led to significant advancements in Natural Language Understanding (NLU). However, PLMs face two main challenges. Firstly, they struggle to generalize to specialized domains like Legal and Finance due to the prevalence of complex domain-specific vocabulary and intricate sentence structures. Secondly, deploying PLMs in real-world applications is difficult due to their high memory requirements and computational demands. To address these issues, we propose Keyword Knowledge Distillation (KKD), a novel approach for in-domain pre-training using selective keyword masking during Knowledge Distillation (KD). KKD transfers knowledge from a domain-specific Teacher BERT model to a smaller, efficient Student BERT model while preserving critical domain-specific information.We evaluate KKD in the Legal and Finance domains across a range of downstream tasks, including multilabel classification, multiclass classification, extractive question answering, regression, multiple-choice question answering, and named entity recognition. The Student models trained using KKD are 40% smaller and 60% faster, while still maintaining high performance compared to the original, domain-specific Teacher models. Specifically, the Student model trained with Legal-Bert preserves 96.4% of the Teacher model’s performance, while the Student model trained with Fin-Bert retains 99.1%. These results underscore KKD’s impressive effectiveness across a variety of tasks and domains.
Momojit Biswas, Anmol Singhal, Preethu Rose Anish
IJCNN3
2024 Generating Clarification Questions for Disambiguating Contracts
abstract
Enterprises frequently enter into commercial contracts that can serve as vital sources of project-specific requirements. Contractual clauses are obligatory, and the requirements derived from contracts can detail the downstream implementation activities that non-legal stakeholders, including requirement analysts, engineers, and delivery personnel, need to conduct. However, comprehending contracts is cognitively demanding and error-prone for such stakeholders due to the extensive use of Legalese and the inherent complexity of contract language. Furthermore, contracts often contain ambiguously worded clauses to ensure comprehensive coverage. In contrast, non-legal stakeholders require a detailed and unambiguous comprehension of contractual clauses to craft actionable requirements. In this work, we introduce a novel legal NLP task that involves generating clarification questions for contracts. These questions aim to identify contract ambiguities on a document level, thereby assisting non-legal stakeholders in obtaining the necessary details for eliciting requirements. This task is challenged by three core issues: (1) data availability, (2) the length and unstructured nature of contracts, and (3) the complexity of legal text. To address these issues, we propose ConRAP, a retrieval-augmented prompting framework for generating clarification questions to disambiguate contractual text. Experiments conducted on contracts sourced from the publicly available CUAD dataset show that ConRAP with ChatGPT can detect ambiguities with an F2 score of 0.87. 70% of the generated clarification questions are deemed useful by human evaluators.
Anmol Singhal, Preethu Rose Anish, Arkajyoti Chakraborty, Smita Ghaisas
LREC/COLING3
2024 Towards Understanding Contracts Grammar: A Large Language Model-Based Extractive Question-Answering Approach
abstract
Software Engineering (SE) contracts play a pivotal role in Information Technology Outsourcing (ITO) projects. The obligations in SE contracts are known to be a useful source for deriving software requirements, thereby contributing to the overall Software Development Life Cycle (SDLC). Making sense of contractual obligations is an important first step in successfully executing software projects. This includes building compliant systems, meeting delivery deadlines, avoiding heavy penalties, and steering clear of expensive litigations. In this work, we present an approach to capture the essence of a contractual clause by extracting its Contracts Grammar. Through an exploratory study, we first identify the constituents of Contracts Grammar. Subsequently, we experiment with multiple approaches for the automated extraction of these constituents, including extractive question-answering, token classification, text-to-text generation, prompting, and regular expressions. The question-answering based approach performed the best in terms of high average ROUGE-L score of 0.81, and faster inference times. The work presented in this paper is a part of the Contracts Governance System (CGS) and is in the process of deployment within a large IT vendor organization.
Gokul Rejithkumar, Preethu Rose Anish, Smita Ghaisas
RE2
2024 Governance-Focused Classification of Security and Privacy Requirements from Obligations in Software Engineering Contracts
Preethu Rose Anish, Aparna Verma, Sivanthy Venkatesan, Logamurugan V., Smita Ghaisas
REFSQ1
2023 A Transformer-based Approach for Abstractive Summarization of Requirements from Obligations in Software Engineering Contracts
abstract
Software Engineering (SE) contracts are a valuable source of software requirements. Seed requirements derived from SE contracts can provide a starting point to the Requirements Engineering (RE) phase. To extract such a seed however, a correct interpretation of contracts text is crucial. A major challenge with contracts text interpretation is that the text is lengthy, convoluted, and it incorporates a complex Legalese. If a summary of the high-level requirements from obligations present in SE contracts is available to the requirement analysts in a language that is comprehensible to them, they can use this seed requirements knowledge to ask the right questions to the stakeholders. In this paper, we propose an approach for summarizing the requirements present in obligations in a language comprehensible to requirement analysts. We use the principles of Prompt Engineering to prompt GPT-3 to generate summaries for training Natural Language Generation (NLG) models for generating SE-specific summaries. Experiments using NLG models such as BART, GPT-2, T5, and Pegasus indicate that Pegasus generates the most accurate summaries with the highest ROUGE score as compared to other models.
Preethu Rose Anish, Smita Ghaisas
RE2
2022 Data is about detail: an empirical investigation for software systems with NLP at core
abstract
Businesses continue to operate under increasingly complex demands such as ever-evolving regulatory landscape, personalization requirements from software apps, and stricter governance with respect to security and privacy. In response to these challenges, large enterprises have been emphasizing automation across a wide range, starting with business processes all the way to customer experience. As AI continues to be a core component of software systems being developed, data assumes a predominant role. AI-centric software systems of industrial scale need large amounts of training data, that in our experience, has introduced several challenges. In this paper, through an empirical study based on interviews with AI practitioners, we present current challenges that need to be addressed in 'data requirements' of Software Systems with NLP at the Core (SSNLPCore). We further discuss the impact of the challenges and techniques currently employed by practitioners for addressing them. Our findings reveal that a focus on details pertaining to data is required early into the project lifecycle, which include aspects such as how we may select, process, and annotate data. This can ensure that the AI component is effective in meeting business goals of software systems.
Anmol Singhal, Preethu Rose Anish, Pratik Sonar, Smita Ghaisas
CAIN2
2021 Analyzing SAFe Practices with Respect to Quality Requirements: Findings from a Qualitative Study
Wasim Alsaqaf, Maya Daneva, Preethu Rose Anish, Roel J. Wieringa
PROFES3
2021 A Pipeline for Automating Labeling to Prediction in Classification of NFRs
abstract
Non-Functional Requirements (NFRs) focus on the operational constraints of the software system. Early detection of NFRs enables their incorporation into the architectural design at an initial stage, a practice obviously preferable to expensive refactoring at a later stage. Automated identification and classification of NFRs has therefore seen numerous efforts using rule-based, machine learning and deep learning-based approaches. One of the major challenges for such an automation is the manual effort that needs to be invested into labeling of training data. This is a concern for large software vendors who typically work on a variety of applications in diverse domains. We address this challenge by designing a pipeline that facilitates classification of NFRs using only a limited amount (~ 20% of an available new dataset) of labeled data for training. We (1) employed Snorkel to automatically label a dataset comprising NFRs from various Software Requirement Specification documents, (2) trained several classifiers using it, and (3) reused these pre-trained classifiers using a Transfer Learning approach to classify NFRs in industry-specific datasets. From among the various language model classifiers, the best results have been obtained for a BERT based classifier fine-tuned to learn the linguistic intricacies of three different domain-specific datasets from real-life projects.
Ranit Chatterjee, Abdul Ahmed, Preethu Rose Anish, Brijendra Suman, Prashant Lawhatre, Smita Ghaisas
RE3
2021 Domain adaptation for an automated classification of deontic modalities in software engineering contracts
abstract
Contracts are agreements between parties engaging in economic transactions. They specify deontic modalities that the signatories should be held responsible for and state the penalties or actions to be taken if the stated agreements are not met. Additionally, contracts have also been known to be source of Software Engineering (SE) requirements. Identifying the deontic modalities in contracts can therefore add value to the Requirements Engineering (RE) phase of SE. The complex and ambiguous language of contracts make it difficult and time-consuming to identify the deontic modalities (obligations, permissions, prohibitions), embedded in the text. State-of-art neural network models are effective for text classification; however, they require substantial amounts of training data. The availability of contracts data is sparse owing to the confidentiality concerns of customers. In this paper, we leverage the linguistic and taxonomical similarities between regulations (available abundantly in the public domain) and contracts to demonstrate that it is possible to use regulations as training data for classifying deontic modalities in real-life contracts. We discuss the results of a range of experiments from the use of rule-based approach to Bidirectional Encoder Representations from Transformers (BERT) for automating the classification of deontic modalities. With BERT, we obtained an average precision and recall of 90% and 89.66% respectively.
Vivek Joshi, Preethu Rose Anish, Smita Ghaisas
ESEC/SIGSOFT FSE2
2020 An Approach That Stimulates Architectural Thinking during Requirements Elicitation: An Empirical Evaluation
abstract
In many global outsourcing projects, the software requirement specifications (SRS) are often orchestrated by requirements analysts who have sufficient business knowledge but are not equipped to ask the kind of questions that are needed to unearth architecturally relevant information from the customer. Often, the resultant SRS therefore lacks some critical details needed by software architects to make informed architectural decisions. To remedy this, the software architects either make assumptions or conduct additional stakeholder interviews resulting in expensive refactoring efforts and project delays. Using an empirical approach, we have designed an approach of using architectural knowledge that can serve as a communication medium between requirements analyst and software architects. In this paper, we present a detailed empirical evaluation of our proposed approach, with practitioners from real-world organizations. Using two studies, we found that in the experience of the participating practitioners, the approach is relevant, easy to use and effective.
Preethu Rose Anish, Maya Daneva, Smita Ghaisas, Roel J. Wieringa
ICSOFT1
2020 Extracting and Classifying Requirements from Software Engineering Contracts
abstract
In this paper, we present our work on extracting and classifying requirements from large software engineering contracts. Typically, the process of requirements elicitation begins after a contractual agreement is signed by all participants. Our interactions with the legal compliance team in a large vendor organization reveal that business contracts can help in the identification of high-level requirements relevant to the success of software engineering projects. We posit that requirements engineering as a discipline has an even wider scope than software engineering of which it is traditionally considered to be a sub-discipline. This is because software engineering-specific requirements are but a part of the success story of any large project. The requirements that emerge from contracts are obligatory in nature, whether or not they pertain to core software development. Therefore, it is important that these are extracted and classified for the benefit of software engineers and other stakeholders responsible for a project. We discuss the results of an exploratory study and a range of experiments from the use of regular expressions to Bidirectional Encoder Representations from Transformers for automating the extraction and classification of requirements from software engineering contracts. With Bidirectional Encoder Representations from Transformers, we obtained a high f-score of greater than eighty four percent for classification of requirements.
Abhishek Sainani, Preethu Rose Anish, Vivek Joshi, Smita Ghaisas
RE2
2016 Probing for requirements knowledge to stimulate architectural thinking
abstract
Software requirements specifications (SRSs) often lack the detail needed to make informed architectural decisions. Architects therefore either make assumptions, which can lead to incorrect decisions, or conduct additional stakeholder interviews, resulting in potential project delays. We previously observed that software architects ask Probing Questions (PQs) to gather information crucial to architectural decision-making. Our goal is to equip Business Analysts with appropriate PQs so that they can ask these questions themselves. We report a new study with over 40 experienced architects to identify reusable PQs for five areas of functionality and organize them into structured flows. These PQ-flows can be used by Business Analysts to elicit and specify architecturally relevant information. Additionally, we leverage machine learning techniques to determine when a PQ-flow is appropriate for use in a project, and to annotate individual PQs with relevant information extracted from the existing SRS. We trained and evaluated our approach on over 8,000 individual requirements from 114 requirements specifications and also conducted a pilot study to validate its usefulness.
Preethu Rose Anish, Balaji Balasubramaniam, Abhishek Sainani, Jane Cleland-Huang, Maya Daneva, Roel J. Wieringa, Smita Ghaisas
ICSE1
2016 Towards an Approach to Stimulate Architectural Thinking during Requirements Engineering Phase
abstract
[Context/Motivation:] The major role of a software architect in the software development life cycle is to design an architectural solution that satisfies the functional and non-functional requirements of the system to the fullest extent possible. However, the details they need to make informed architectural decisions are often missing from the software requirements specifications. Architects therefore either make assumptions or go back to the business analysts for clarifications or conduct additional stakeholder interviews. All of these result in potential project delays. [Question/Problem:] There is very little known about how architects gather this missing information that is crucial to architectural decision-making. Further, there is dearth of knowledge on how we can leverage the requirements analysis phase by empowering business analysts to gather architecturally significant information pertinent to a given requirement. [Principal ideas / Results:] Based on qualitative interview study with software architects from large organisations, this research will (a) investigate how architects cope with missing information in real-world large scale projects and (b) leverage the knowledge of experienced software architects and make it available to business analysts so that they are equipped to elicit a more complete set of requirements that feed sufficient information into the architecture design process. This research further explores the use of machine learning techniques to gauge the potential to introduce automation to the process of gathering the architecturally significant information missing in the software requirement specification.
Preethu Rose Anish
RE1
2015 What you ask is what you get: Understanding architecturally significant functional requirements
abstract
Software architects are responsible for designing an architectural solution that satisfies the functional and non-functional requirements of the system to the fullest extent possible. However, the details they need to make informed architectural decisions are often missing from the requirements specification. An earlier study we conducted indicated that architects intuitively recognize architecturally significant requirements in a project, and often seek out relevant stakeholders in order to ask Probing Questions (PQs) that help them acquire the information they need. This paper presents results from a qualitative interview study aimed at identifying architecturally significant functional requirements' categories from various business domains, exploring relevant PQs for each category, and then grouping PQs by type. Using interview data from 14 software architects in three countries, we identified 15 categories of architecturally significant functional requirements and 6 types of PQs. We found that the domain knowledge of the architect and her experience influence the choice of PQs significantly. A preliminary quantitative evaluation of the results against real-life software requirements specification documents indicated that software specifications in our sample largely do not contain the crucial architectural differentiators that may impact architectural choices and that PQs are a necessary mechanism to unearth them. Further, our findings provide the initial list of PQs which could be used to prompt business analysts to elicit architecturally significant functional requirements that the architects need.
Preethu Rose Anish, Maya Daneva, Jane Cleland-Huang, Roel J. Wieringa, Smita Ghaisas
RE1
2014 Product knowledge configurator for requirements gap analysis and customizations
abstract
Product knowledge plays an important role in identifying the requirements of the desired variant and configuring the existing product to the present needs of a customer. The success of a product-based business depends to a great extent on how efficiently and accurately the existing product knowledge is utilized for customization needs. Oftentimes however, product knowledge resides with few key individuals in an organization. In the absence of their involvement, project teams may redevelop product features unnecessarily, resulting in an effort overhead. Such overdependence poses a risk to projects. To identify the requirements for the variants accurately and efficiently, we need to have a thorough knowledge of the existing product features. In this paper, we discuss our work on representing product knowledge and reusing it in a Requirements Engineering (RE) exercise for a large project involving product customization. We present our experience from using the configurator for requirements gap analysis and customizations.
Preethu Rose Anish, Smita Ghaisas
RE1
2013 Detecting system use cases and validations from documents
abstract
Identifying system use cases and corresponding validations involves analyzing large requirement documents to understand the descriptions of business processes, rules and policies. This consumes a significant amount of effort and time. We discuss an approach to automate the detection of system use cases and corresponding validations from documents. We have devised a representation that allows for capturing the essence of rule statements as a composition of atomic `Rule intents' and key phrases associated with the intents. Rule intents that co-occur frequently constitute `Rule acts' analogous to the Speech acts in Linguistics. Our approach is based on NLP techniques designed around this Rule Model. We employ syntactic and semantic NL analyses around the model to identify and classify rules and annotate them with Rule acts. We map the Rule acts to business process steps and highlight the combinations as potential system use cases and validations for human supervision.
Smita Ghaisas, Manish Motwani, Preethu Rose Anish
ASE3