VLDB 2026 Research / reviewers in the wild / expert
Mitra Bokaei Hosseini
dblp:179/8585
· DBLP profile ↗
10ranked-venue papers
5as first author
5since 2021 · last 2025
0000-0001-8069-0088ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 10 · 5 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Demystifying Feature Requests: Leveraging LLMs to Refine Feature Requests in Open-Source SoftwareabstractThe growing popularity and widespread use of software applications (apps) across various domains have driven rapid industry growth. Along with this growth, fast-paced market changes have led to constantly evolving software requirements. Such requirements are often grounded in feature requests and enhancement suggestions, typically provided by users in natural language (NL). However, these requests often suffer from defects such as ambiguity and incompleteness, making them challenging to interpret. Traditional validation methods (e.g., interviews and workshops) help clarify such defects but are impractical in decentralized environments like open-source software (OSS), where change requests originate from diverse users on platforms like GitHub. This paper proposes a novel approach leveraging Large Language Models (LLMs) to detect and refine NL defects in feature requests. Our approach automates the identification of ambiguous and incomplete requests and generates clarification questions (CQs) to enhance their usefulness for developers. To evaluate its effectiveness, we apply our method to real-world OSS feature requests and compare its performance against human annotations. In addition, we conduct interviews with GitHub developers to gain deeper insights into their perceptions of NL defects, the strategies they use to address these defects, and the impact of defects on downstream software engineering (SE) tasks. Pragyan K. C, Rambod Ghandiparsi, Thomas Herron, John Heaps, Mitra Bokaei Hosseini |
RE | 5 |
| 2024 | A Large Language Model Approach to Code and Privacy Policy AlignmentabstractAs mobile technology has advanced, individuals have started relying on their smartphones to conduct more of their everyday tasks. From playing games or streaming media to social networking and banking, apps on a user's device may have access to the most sensitive information on the device. Privacy policies are designed to inform users of such data practices so that they can make reasonable decisions when using the app. However, an app's true behavior may not always align with the statements in a privacy policy. In this work, we divide our study into two components and compare the viability of various large language models (LLMs): methods for extracting and summarizing privacy policy data practices, or information-type extraction and action-verb extraction; and methods for measuring whether the policy acknowledges the interaction with certain information (sensitive data) compared to identified methods within its app's source code. Fine-tuning GPT-3.5 Turbo delivers a higher average F1-score for both action verb extraction (0.50) and information-type extraction (0.84) compared to other LLMs. ChatGPT outperforms other language models in traditional semantic similarity, providing a consistently high performance, including the highest F1-score (0.52) for this task. Our approaches demonstrate that these LLMs are viable in performing such tasks and additionally that pre-trained instruction-based LLMs are capable of identifying the complex relationships between policies and source code. Gabriel A. Morales, Pragyan K. C, Sadia Jahan, Mitra Bokaei Hosseini, Rocky Slavin |
SANER | 4 |
| 2023 | Mobile Application Privacy Risk Assessments from User-authored ScenariosabstractMobile applications (apps) provide users valuable benefits at the risk of exposing users to privacy harms. Improving privacy in mobile apps faces several challenges, in particular, that many apps are developed by low resourced software development teams, such as end-user programmers or in startups. In addition, privacy risks are primarily known to users, which can make it difficult for developers to prioritize privacy for sensitive data. In this paper, we introduce a novel, lightweight method that allows app developers to elicit scenarios and privacy risk scores from users directly using only an app screenshot. The technique relies on named entity recognition (NER) to identify information types in user-authored scenarios, which are then fed in real-time to a privacy risk survey that users complete. The best-performing NER model predicts information types with a weighted average precision of 0.70 and recall of 0.72, after post-processing to remove false positives. The model was trained on a labeled 300-scenario corpus, and evaluated in an end-to-end evaluation using an additional 203 scenarios yielding 2,338 user-provided privacy risk scores. Finally, we discuss how developers can use the risk scores to prioritize, select and apply privacy design strategies in the context of four user-authored scenarios. Tianjian Huang, Vaishnavi Kaulagi, Mitra Bokaei Hosseini, Travis D. Breaux |
RE | 3 |
| 2021 | Ambiguity and Generality in Natural Language Privacy PoliciesabstractPrivacy policies are legal documents containing application data practices. These documents are well-established sources of requirements in software engineering. However, privacy policies are written in natural language, thus subject to ambiguity and abstraction. Eliciting requirements from privacy policies is a challenging task as these ambiguities can result in more than one interpretation of a given information type (e.g., ambiguous information type "device information" in the statement "we collect your device information"). To address this challenge, we propose an automated approach to infer semantic relations among information types and construct an ontology to guide requirements authors in the selection of the most appropriate information type terms. Our solution utilizes word embeddings and Convolutional Neural Networks (CNN) to classify information type pairs as either hypernymy, synonymy, or unknown. We evaluate our model on a manually-built ontology, yielding predictions that identify hypernymy relations in information type pairs with 0.904 F-1 score, suggesting a large reduction in effort required for ontology construction. Mitra Bokaei Hosseini, John Heaps, Rocky Slavin, Jianwei Niu 0001, Travis D. Breaux |
RE | 1 |
| 2021 | Analyzing privacy policies through syntax-driven semantic analysis of information types
Mitra Bokaei Hosseini, Travis D. Breaux, Rocky Slavin, Jianwei Niu 0001, Xiaoyin Wang |
Inf. Softw. Technol. | 1 |
| 2020 | Disambiguating Requirements Through Syntax-Driven Semantic Analysis of Information Types
Mitra Bokaei Hosseini, Rocky Slavin, Travis D. Breaux, Xiaoyin Wang, Jianwei Niu 0001 |
REFSQ | 1 |
| 2018 | GUILeak: tracing privacy policy claims on user input data for Android applicationsabstractThe Android mobile platform supports billions of devices across more than 190 countries around the world. This popularity coupled with user data collection by Android apps has made privacy protection a well-known challenge in the Android ecosystem. In practice, app producers provide privacy policies disclosing what information is collected and processed by the app. However, it is difficult to trace such claims to the corresponding app code to verify whether the implementation is consistent with the policy. Existing approaches for privacy policy alignment focus on information directly accessed through the Android platform (e.g., location and device ID), but are unable to handle user input, a major source of private information. In this paper, we propose a novel approach that automatically detects privacy leaks of user-entered data for a given Android app and determines whether such leakage may violate the app's privacy policy claims. For evaluation, we applied our approach to 120 popular apps from three privacy-relevant app categories: finance, health, and dating. The results show that our approach was able to detect 21 strong violations and 18 weak violations from the studied apps. Xiaoyin Wang, Mitra Bokaei Hosseini, Rocky Slavin, Travis D. Breaux, Jianwei Niu 0001 |
ICSE | 3 |
| 2018 | Inferring Ontology Fragments from Semantic Role Typing of Lexical Variants
Mitra Bokaei Hosseini, Travis D. Breaux, Jianwei Niu 0001 |
REFSQ | 1 |
| 2018 | Semantic inference from natural language privacy policies and Android codeabstractMobile apps collect dierent categories of personal information to provide users with various services. Companies use privacy policies containing critical requirements to inform users about their data practices. With the growing access to personal information and the scale of mobile app deployment, traceability of links between privacy policy requirements and app code is increasingly important. Automated traceability can be achieved using natural language processing and code analysis techniques. However, such techniques must address two main challenges: ambiguity in privacy policy terminology and unbounded information types provided by users through input elds in GUI. In this work, we propose approaches to interpret abstract terms in privacy policies, identify information types in Android layout code, and create a mapping between them using natural language processing techniques. Mitra Bokaei Hosseini |
ESEC/SIGSOFT FSE | 1 |
| 2016 | Toward a framework for detecting privacy policy violations in android application codeabstractMobile applications frequently access sensitive personal information to meet user or business requirements. Because such information is sensitive in general, regulators increasingly require mobile-app developers to publish privacy policies that describe what information is collected. Furthermore, regulators have fined companies when these policies are inconsistent with the actual data practices of mobile apps. To help mobile-app developers check their privacy policies against their apps' code for consistency, we propose a semi-automated framework that consists of a policy terminology-API method map that links policy phrases to API methods that produce sensitive information, and information flow analysis to detect misalignments. We present an implementation of our framework based on a privacy-policy-phrase ontology and a collection of mappings from API methods to policy phrases. Our empirical evaluation on 477 top Android apps discovered 341 potential privacy policy violations. Rocky Slavin, Xiaoyin Wang, Mitra Bokaei Hosseini, James Hester, Ram Krishnan, Jaspreet Bhatia, Travis D. Breaux, Jianwei Niu 0001 |
ICSE | 3 |