Vijayanta Jain

dblp:250/9339 · DBLP profile ↗
← Back
6ranked-venue papers
4as first author
6since 2021 · last 2023
0000-0003-2652-5107ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Security and privacy · 2 · 2 first-author · 2 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2023 Towards Fine-Grained Localization of Privacy Behaviors
abstract
Privacy labels help developers communicate their application’s privacy behaviors (i.e., how and why an application uses personal information) to users. But, studies show that developers face several challenges in creating them and the resultant labels are often inconsistent with their application’s privacy behaviors. In this paper, we create a novel methodology called fine-grained localization of privacy behaviors to locate individual statements in source code which encode privacy behaviors and predict their privacy labels. We design and develop an attention-based multi-head encoder model which creates individual representations of multiple methods and uses attention to identify relevant statements that implement privacy behaviors. These statements are then used to predict privacy labels for the application’s source code and can help developers write privacy statements that can be used as notices. Our quantitative analysis shows that our approach can achieve high accuracy in identifying privacy labels, with the lowest accuracy of 91.41% and the highest of 98.45%. We also evaluate the efficacy of our approach with six software professionals from our university. The results demonstrate that our approach reduces the time and mental effort required by developers to create high-quality privacy statements and can finely localize statements in methods that implement privacy behaviors.
Vijayanta Jain, Sepideh Ghanavati, Sai Teja Peddinti, Collin McMillan
EuroS&P1
2023 A Language Model of Java Methods with Train/Test Deduplication
abstract
This tool demonstration presents a research toolkit for a language model of Java source code. The target audience includes researchers studying problems at the granularity level of subroutines, statements, or variables in Java. In contrast to many existing language models, we prioritize features for researchers including an open and easily-searchable training set, a held out test set with different levels of deduplication from the training set, infrastructure for deduplicating new examples, and an implementation platform suitable for execution on equipment accessible to a relatively modest budget. Our model is a GPT2-like architecture with 350m parameters. Our training set includes 52m Java methods (9b tokens) and 13m StackOverflow threads (10.5b tokens). To improve accessibility of research to more members of the community, we limit local resource requirements to GPUs with 16GB video memory. We provide a test set of held out Java methods that include descriptive comments, including the entire Java projects for those methods. We also provide deduplication tools using precomputed hash tables at various similarity thresholds to help researchers ensure that their own test examples are not in the training set. We make all our tools and data open source and available via Huggingface and Github.
Chia-Yi Su, Aakash Bansal, Vijayanta Jain, Sepideh Ghanavati, Collin McMillan
ESEC/SIGSOFT FSE3
2022 Creating Consistent Privacy Notices by Translating Code Segments into Privacy Captions
abstract
A privacy notice is a short descriptive statement that describes to users how their information is used and why?. When a mobile application uses any personal information of the user, such as location, it must provide them with a privacy notice. However, creating concise and accurate privacy notices with each update to an application is a challenging task. Previous efforts have focused on creating these notices through questionnaires or predefined templates which do not make application source code traceable with privacy notices. Lack of traceability between source code and privacy notices leads to inconsistencies and inaccuracies. In this paper, we discuss our approach to creating privacy captions, short sentences that describe how and why users’ personal information is used. by translating the application’s source code. In this work, we explain our plan to implement our approach and demonstrate the steps we have implemented so far.
Vijayanta Jain
RE1
2022 PAcT: Detecting and Classifying Privacy Behavior of Android Applications
abstract
Interpreting and describing mobile applications' privacy behaviors to ensure creating consistent and accurate privacy notices is a challenging task for developers. Traditional approaches to creating privacy notices are based on predefined templates or questionnaires and do not rely on any traceable behaviors in code which may result in inconsistent and inaccurate notices. In this paper, we present an automated approach to detect privacy behaviors in code of Android applications. We develop Privacy Action Taxonomy (PAcT), which includes labels for Practice (i.e. how applications use personal information) and Purpose (i.e. why). We annotate ~5,200 code segments based on the labels and create a multi-label multi-class dataset with ~14,000 labels. We develop and train deep learning models to classify code segments. We achieve the highest F-1 scores across all label types of 79.62% and 79.02% for Practice and Purpose.
Vijayanta Jain, Sanonda Datta Gupta, Sepideh Ghanavati, Sai Teja Peddinti, Collin McMillan
WISEC1
2021 PHIN: A Privacy Protected Heterogeneous IoT Network
Sanonda Datta Gupta, Aubree Nygaard, Stephen Kaplan, Vijayanta Jain, Sepideh Ghanavati
RCIS4
2021 PriGen: Towards Automated Translation of Android Applications' Code to Privacy Captions
Vijayanta Jain, Sanonda Datta Gupta, Sepideh Ghanavati, Sai Teja Peddinti
RCIS1