VLDB 2026 Research / reviewers in the wild / expert
Thomas B. Norton
dblp:183/8184
· DBLP profile ↗
10ranked-venue papers
0as first author
6since 2021 · last 2024
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 3 since 2021Software engineering, systems software and programming languages · 3 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Creation and Analysis of an International Corpus of Privacy LawsabstractThe landscape of privacy laws and regulations around the world is complex and ever-changing. National and super-national laws, agreements, decrees, and other government-issued rules form a patchwork that companies must follow to operate internationally. To examine the status and evolution of this patchwork, we introduce the Privacy Law Corpus, of 1,043 privacy laws, regulations, and guidelines, covering 183 jurisdictions. This corpus enables a large-scale quantitative and qualitative examination of legal focus on privacy. We examine the temporal distribution of when privacy laws were created and illustrate the dramatic increase in privacy legislation over the past 50 years, although a finer-grained examination reveals that the rate of increase varies depending on the personal data types that privacy laws address. Our exploration also demonstrates that most privacy laws respectively address relatively few personal data types. Additionally, topic modeling results show the prevalence of common themes in privacy laws, such as finance, healthcare, and telecommunications. Finally, we release the corpus to the research community to promote further study. Sonu Gupta, Geetika Gopi, Harish Balaji, Ellen Poplavska, Nora O'Toole, Siddhant Arora, Thomas B. Norton, Norman M. Sadeh, Shomir Wilson |
LREC/COLING | 7 |
| 2024 | Requirements Satisfiability with In-Context LearningabstractLanguage models that can learn a task at inference time, called in-context learning (ICL), show increasing promise in natural language inference tasks. In ICL, a model user constructs a prompt to describe a task with a natural language instruction and zero or more examples, called demonstrations. The prompt is then input to the language model to generate a completion. In this paper, we apply ICL to the design and evaluation of satisfaction arguments, which describe how a requirement is satisfied by a system specification and associated domain knowledge. The approach builds on three prompt design patterns, including augmented generation, prompt tuning, and chain-of-thought prompting, and is evaluated on a privacy problem to check whether a mobile app scenario and associated design description satisfies eight consent requirements from the EU General Data Protection Regulation (GDPR). The overall results show that GPT-4 can be used to verify requirements satisfaction with 96.7% accuracy and dissatisfaction with 93.2% accuracy. Inverting the requirement improves verification of dissatisfaction to 97.2%. Chain-of-thought prompting improves overall GPT-3.5 performance by 9.0% accuracy. We discuss the trade-offs among templates, models and prompt strategies and provide a detailed analysis of the generated specifications to inform how the approach can be applied in practice. Sarah Santos, Travis D. Breaux, Thomas B. Norton, Sara Haghighi, Sepideh Ghanavati |
RE | 3 |
| 2024 | Incorporating Taxonomic Reasoning and Regulatory Knowledge into Automated Privacy Question Answering
Abhilasha Ravichander, Ian Yang, Rex Chen, Shomir Wilson, Thomas B. Norton, Norman M. Sadeh |
WISE (1) | 5 |
| 2022 | A Tale of Two Regulatory Regimes: Creation and Analysis of a Bilingual Privacy Policy CorpusabstractOver the past decade, researchers have started to explore the use of NLP to develop tools aimed at helping the public, vendors, and regulators analyze disclosures made in privacy policies. With the introduction of new privacy regulations, the language of privacy policies is also evolving, and disclosures made by the same organization are not always the same in different languages, especially when used to communicate with users who fall under different jurisdictions. This work explores the use of language technologies to capture and analyze these differences at scale. We introduce an annotation scheme designed to capture the nuances of two new landmark privacy regulations, namely the EU’s GDPR and California’s CCPA/CPRA. We then introduce the first bilingual corpus of mobile app privacy policies consisting of 64 privacy policies in English (292K words) and 91 privacy policies in German (478K words), respectively with manual annotations for 8K and 19K fine-grained data practices. The annotations are used to develop computational methods that can automatically extract “disclosures” from privacy policies. Analysis of a subset of 59 “semi-parallel” policies reveals differences that can be attributed to different regulatory regimes, suggesting that systematic analysis of policies using automated language technologies is indeed a worthwhile endeavor. Siddhant Arora, Henry Hosseini, Christine Utz, Vinayshekhar Bannihatti Kumar, Tristan Dhellemmes, Abhilasha Ravichander, Peter Story, Jasmine Mangat, Rex Chen, Martin Degeling, Thomas B. Norton, Thomas Hupperich, Shomir Wilson, Norman M. Sadeh |
LREC | 11 |
| 2022 | Legal Accountability as Software Quality: A U.S. Data Processing PerspectiveabstractSoftware and hardware innovation has led to new consumer products and services with significant benefits to consumers and society. These advances, however, can come with great cost to society when they fail to comply with government laws and regulations. While compliance failures do result from technical missteps in design, there is also a wide gap between the technical expertise and culture of legal analysts and software engineers, as well as competing priorities between legal requirements and business objectives. In this perspective paper, we propose changing legal compliance from a corporate oversight activity to a principal design activity, wherein lawyers and software engineers employ enhanced methods and tools tailored to bridge the cultural and knowledge gap and assess legal and business trade-offs. To that end, we describe a new software quality, called Legal Accountability, which can be evaluated alongside other qualities, such as usability, modifiability, performance and testing. Legal Accountability has five properties that lawyers and designers must attend to, including legal traceability, completeness, validity, auditability and continuity. We illustrate the quality with examples from the U.S. data processing perspective, and prior work in requirements engineering, before concluding with future and ongoing research challenges. Travis D. Breaux, Thomas B. Norton |
RE | 2 |
| 2021 | Breaking Down Walls of Text: How Can NLP Benefit Consumer Privacy?abstractAbhilasha Ravichander, Alan W Black, Thomas Norton, Shomir Wilson, Norman Sadeh. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Abhilasha Ravichander, Alan W. Black, Thomas B. Norton, Shomir Wilson, Norman M. Sadeh |
ACL/IJCNLP (1) | 3 |
| 2020 | From Prescription to Description: Mapping the GDPR to a Privacy Policy Corpus Annotation SchemeabstractThe European Union’s General Data Protection Regulation (GDPR) has compelled businesses and other organizations to update their privacy policies to state specific information about their data practices. Simultaneously, researchers in natural language processing (NLP) have developed corpora and annotation schemes for extracting salient information from privacy policies, often independently of specific laws. To connect existing NLP research on privacy policies with the GDPR, we introduce a mapping from GDPR provisions to the OPP-115 annotation scheme, which serves as the basis for a growing number of projects to automatically classify privacy policy text. We show that assumptions made in the annotation scheme about the essential topics for a privacy policy reflect many of the same topics that the GDPR requires in these documents. This suggests that OPP-115 continues to be representative of the anatomy of a legally compliant privacy policy, and that the legal assumptions behind it represent the elements of data processing that ought to be disclosed within a policy for transparency. The correspondences we show between OPP-115 and the GDPR suggest the feasibility of bridging existing computational and legal research on privacy policies, benefiting both areas. Ellen Poplavska, Thomas B. Norton, Shomir Wilson, Norman M. Sadeh |
JURIX | 2 |
| 2019 | Question Answering for Privacy Policies: Combining Computational and Legal PerspectivesabstractAbhilasha Ravichander, Alan W Black, Shomir Wilson, Thomas Norton, Norman Sadeh. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Abhilasha Ravichander, Alan W. Black, Shomir Wilson, Thomas B. Norton, Norman M. Sadeh |
EMNLP/IJCNLP (1) | 4 |
| 2016 | The Creation and Analysis of a Website Privacy Policy CorpusabstractShomir Wilson, Florian Schaub, Aswarth Abhilash Dara, Frederick Liu, Sushain Cherivirala, Pedro Giovanni Leon, Mads Schaarup Andersen, Sebastian Zimmeck, Kanthashree Mysore Sathyendra, N. Cameron Russell, Thomas B. Norton, Eduard Hovy, Joel Reidenberg, Norman Sadeh. Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2016. Shomir Wilson, Florian Schaub, Aswarth Abhilash Dara, Frederick Liu, Sushain Cherivirala, Pedro Giovanni Leon, Mads Schaarup Andersen, Sebastian Zimmeck, Kanthashree Mysore Sathyendra, N. Cameron Russell, Thomas B. Norton, Eduard H. Hovy, Joel R. Reidenberg, Norman M. Sadeh |
ACL (1) | 11 |
| 2016 | A Theory of Vagueness and Privacy Risk PerceptionabstractAmbiguity arises in requirements when a statement is unintentionally or otherwise incomplete, missing information, or when a word or phrase has more than one possible meaning. For web-based and mobile information systems, ambiguity, and vagueness inparticular, undermines the ability of organizations to align their privacy policies with their data practices, which can confuse or mislead users thus leading to an increase in privacy risk. In this paper, we introduce a theory of vagueness for privacy policy statements based on a taxonomy of vague terms derived from an empirical content analysis of 15 privacy policies. The taxonomy was evaluated in a paired comparison experiment and results were analyzed using the Bradley-Terry model to yield a rank order of vague terms in both isolation and composition. The theory predicts how vague modifiers to information actions and information types can be composed to increase or decrease overall vagueness. We further provide empirical evidence based on factorial vignette surveys to show how increases in vagueness will decrease users' acceptance of privacy risk and thus decrease users' willingness to share personal information. Jaspreet Bhatia, Travis D. Breaux, Joel R. Reidenberg, Thomas B. Norton |
RE | 4 |