Anmol Singhal

dblp:297/5503 · DBLP profile ↗
← Back
6ranked-venue papers
3as first author
6since 2021 · last 2025
0009-0003-9122-4789ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 2 first-author · 4 since 2021Software engineering, systems software and programming languages · 3 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 From Legal Text to Tech Specs: Generative AI's Interpretation of Consent in Privacy Law
abstract
Privacy law and regulation have turned to “consent” as the legitimate basis for collecting and processing individuals’ data. As governments have rushed to enshrine consent requirements in their privacy laws, such as the California Consumer Privacy Act (CCPA), significant challenges remain in understanding how these legal mandates are operationalized in software. The opaque nature of software development processes further complicates this translation. To address this, we explore the use of Large Language Models (LLMs) in requirements engineering to bridge the gap between legal requirements and technical implementation. This study employs a three-step pipeline that involves using an LLM to classify software use cases for compliance, generating LLM modifications for non-compliant cases, and manually validating these changes against legal standards. Our preliminary findings highlight the potential of LLMs in automating compliance tasks, while also revealing limitations in their reasoning capabilities. By benchmarking LLMs against real-world use cases, this research provides insights into leveraging AI-driven solutions to enhance legal compliance of software.
Aniket Kesari, Travis D. Breaux, Sarah Santos, Tom Norton, Anmol Singhal
ICAIL5
2025 Optimized Domain-Specific Text Processing with Keyword Knowledge Distillation (KKD)
abstract
The generative and reasoning capabilities of Pre-Trained Language Models (PLMs) have led to significant advancements in Natural Language Understanding (NLU). However, PLMs face two main challenges. Firstly, they struggle to generalize to specialized domains like Legal and Finance due to the prevalence of complex domain-specific vocabulary and intricate sentence structures. Secondly, deploying PLMs in real-world applications is difficult due to their high memory requirements and computational demands. To address these issues, we propose Keyword Knowledge Distillation (KKD), a novel approach for in-domain pre-training using selective keyword masking during Knowledge Distillation (KD). KKD transfers knowledge from a domain-specific Teacher BERT model to a smaller, efficient Student BERT model while preserving critical domain-specific information.We evaluate KKD in the Legal and Finance domains across a range of downstream tasks, including multilabel classification, multiclass classification, extractive question answering, regression, multiple-choice question answering, and named entity recognition. The Student models trained using KKD are 40% smaller and 60% faster, while still maintaining high performance compared to the original, domain-specific Teacher models. Specifically, the Student model trained with Legal-Bert preserves 96.4% of the Teacher model’s performance, while the Student model trained with Fin-Bert retains 99.1%. These results underscore KKD’s impressive effectiveness across a variety of tasks and domains.
Momojit Biswas, Anmol Singhal, Preethu Rose Anish
IJCNN2
2025 Requirements Elicitation Follow-Up Question Generation
abstract
Interviews are a widely used technique in eliciting requirements to gather stakeholder needs, preferences, and expectations for a software system. Effective interviewing requires skilled interviewers to formulate appropriate interview questions in real time while facing multiple challenges, including lack of familiarity with the domain, excessive cognitive load, and information overload that hinders how humans process stakeholders’ speech. Recently, large language models (LLMs) have exhibited state-of-the-art performance in multiple natural language processing tasks, including text summarization and entailment. To support interviewers, we investigate the application of GPT-4o to generate follow-up interview questions during requirements elicitation by building on a framework of common interviewer mistake types. In addition, we describe methods to generate questions based on interviewee speech. We report a controlled experiment to evaluate LLM-generated and human-authored questions with minimal guidance, and a second controlled experiment to evaluate the LLM-generated questions when generation is guided by interviewer mistake types. Our findings demonstrate that, for both experiments, the LLM-generated questions are no worse than the human-authored questions with respect to clarity, relevancy, and informativeness. In addition, LLM-generated questions outperform human-authored questions when guided by common mistakes types. This highlights the potential of using LLMs to help interviewers improve the quality and ease of requirements elicitation interviews in real time.
Anmol Singhal, Travis D. Breaux
RE2
2025 Legal Requirements Translation from Law
abstract
Software systems must comply with legal regulations, which is a resource-intensive task, particularly for small organizations and startups lacking dedicated legal expertise. Extracting metadata from regulations to elicit legal requirements for software is a critical step to ensure compliance. However, it is a cumbersome task due to the length and complex nature of legal text. Although prior work has pursued automated methods for extracting structural and semantic metadata from legal text, key limitations remain: they do not consider the interplay and interrelationships among attributes associated with these metadata types, and they rely on manual labeling or heuristic-driven machine learning, which does not generalize well to new documents.In this paper, we introduce an approach based on textual entailment and in-context learning for automatically generating a canonical representation of legal text—encodable and executable as Python code. Our representation is instantiated from a manually designed Python class structure that serves as a domain-specific metamodel, capturing both structural and semantic legal metadata and their interrelationships. This design choice reduces the need for large, manually labeled datasets and enhances applicability to unseen legislation. We evaluate our approach on 13 U.S. state data breach notification laws, demonstrating that our generated representations pass approximately 89.4% of test cases and achieve a precision and recall of 82.2 and 88.7 respectively.
Anmol Singhal, Travis D. Breaux
RE1
2024 Generating Clarification Questions for Disambiguating Contracts
abstract
Enterprises frequently enter into commercial contracts that can serve as vital sources of project-specific requirements. Contractual clauses are obligatory, and the requirements derived from contracts can detail the downstream implementation activities that non-legal stakeholders, including requirement analysts, engineers, and delivery personnel, need to conduct. However, comprehending contracts is cognitively demanding and error-prone for such stakeholders due to the extensive use of Legalese and the inherent complexity of contract language. Furthermore, contracts often contain ambiguously worded clauses to ensure comprehensive coverage. In contrast, non-legal stakeholders require a detailed and unambiguous comprehension of contractual clauses to craft actionable requirements. In this work, we introduce a novel legal NLP task that involves generating clarification questions for contracts. These questions aim to identify contract ambiguities on a document level, thereby assisting non-legal stakeholders in obtaining the necessary details for eliciting requirements. This task is challenged by three core issues: (1) data availability, (2) the length and unstructured nature of contracts, and (3) the complexity of legal text. To address these issues, we propose ConRAP, a retrieval-augmented prompting framework for generating clarification questions to disambiguate contractual text. Experiments conducted on contracts sourced from the publicly available CUAD dataset show that ConRAP with ChatGPT can detect ambiguities with an F2 score of 0.87. 70% of the generated clarification questions are deemed useful by human evaluators.
Anmol Singhal, Preethu Rose Anish, Arkajyoti Chakraborty, Smita Ghaisas
LREC/COLING1
2022 Data is about detail: an empirical investigation for software systems with NLP at core
abstract
Businesses continue to operate under increasingly complex demands such as ever-evolving regulatory landscape, personalization requirements from software apps, and stricter governance with respect to security and privacy. In response to these challenges, large enterprises have been emphasizing automation across a wide range, starting with business processes all the way to customer experience. As AI continues to be a core component of software systems being developed, data assumes a predominant role. AI-centric software systems of industrial scale need large amounts of training data, that in our experience, has introduced several challenges. In this paper, through an empirical study based on interviews with AI practitioners, we present current challenges that need to be addressed in 'data requirements' of Software Systems with NLP at the Core (SSNLPCore). We further discuss the impact of the challenges and techniques currently employed by practitioners for addressing them. Our findings reveal that a focus on details pertaining to data is required early into the project lifecycle, which include aspects such as how we may select, process, and annotate data. This can ensure that the AI component is effective in meeting business goals of software systems.
Anmol Singhal, Preethu Rose Anish, Pratik Sonar, Smita Ghaisas
CAIN1