EDBT 2026 Demo / reviewers in the wild / expert
Razieh Nokhbeh Zaeem
dblp:99/7855
· DBLP profile ↗
20ranked-venue papers
12as first author
3since 2021 · last 2022
0000-0002-0415-5814ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 8 · 5 first-authorArtificial intelligence and machine learning · 6 · 4 first-author · 2 since 2021Security and privacy · 4 · 2 first-author · 1 since 2021Databases, data management, data science and information retrieval · 4 · 3 first-author · 1 since 2021Computer networks · 2 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Software engineering, system software, and programming languages
2 papers |
Software testing · 57% Debugging and program repair · 38% Program analysis · 6% | |
| Network and information security
1 paper |
Privacy and data protection · 100% |
Topics — the 7 heaviest of 7, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Privacy and data protection › privacy policy
privacy policy analysis |
0.6 | 1 | 2022 | PrivacyCheck v3: Empowering Users with Higher-Level Understanding of Privacy Policies · WSDM 2022 |
Software testing › test generation
constraint-based test generation |
0.1 | 1 | 2012 | Test input generation using dynamic programming · SIGSOFT FSE 2012 |
Debugging and program repair
fault localization |
0.1 | 1 | 2012 | Improving the effectiveness of spectra-based fault localization using specifications · ASE 2012 |
Software testing › test generation
random test generation |
0.1 | 1 | 2012 | Test input generation using dynamic programming · SIGSOFT FSE 2012 |
Debugging and program repair › fault localization
spectrum-based fault localization |
0.1 | 1 | 2012 | Improving the effectiveness of spectra-based fault localization using specifications · ASE 2012 |
Software testing
test input generation |
0.1 | 1 | 2012 | Test input generation using dynamic programming · SIGSOFT FSE 2012 |
Program analysis
symbolic execution |
0.0 | 1 | 2012 | Test input generation using dynamic programming · SIGSOFT FSE 2012 |
Methods — techniques the papers use, named apart from their topics
traffic analysis · 0.6machine learning · 0.6unsatisfiable core computation · 0.1symbolic execution · 0.1lazy initialization · 0.1dynamic programming · 0.1SAT solving · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | PrivacyCheck v3: Empowering Users with Higher-Level Understanding of Privacy PoliciesabstractOnline privacy policies are lengthy and hard to read, yet are profoundly important as they communicate the practices of an organization pertaining to user data privacy. Privacy Enhancing Technologies, or PETs, seek to inform users by summarizing these privacy policies. Efforts in the research and development of such PETs, however, have largely been limited to tools that recap the policy or visualize it. We present the next generation of our research and publicly available tool, PrivacyCheck v3, that utilizes machine learning to inform and empower users with respect to privacy policies. PrivacyCheck v3 adds capabilities that are commonly absent from similar PETs on the web. In particular, it adds the ability to (1) find the competitors of an organization with Alexa traffic analysis and compare policies across them, (2) follow privacy policies to which the user has agreed and notify the user when policies change, (3) track policies over time and report how often policies change and their trends, (4) automatically find privacy policies in domains, and (5) provide a bird's-eye view of privacy policies. The new features of PrivacyCheck not only inform users about details of privacy policies, but also empower them to understand privacy policies at a higher level, make informed decisions, and even select competitors with better privacy policies. Razieh Nokhbeh Zaeem, Ahmad Ahbab, Josh Bestor, Hussam H. Djadi, Sunny Kharel, Victor Lai, Nick Wang, K. Suzanne Barber |
WSDM | 1 |
| 2021 | A Large Publicly Available Corpus of Website Privacy Policies Based on DMOZabstractStudies have shown website privacy policies are too long and hard to comprehend for their target audience. These studies and a more recent body of research that utilizes machine learning and natural language processing to automatically summarize privacy policies greatly benefit, if not rely on, corpora of privacy policies collected from the web. While there have been smaller annotated corpora of web privacy policies made public, we are not aware of any large publicly available corpus. We use DMOZ, a massive open-content directory of the web, and its manually categorized 1.5 million websites, to collect hundreds of thousands of privacy policies associated with their categories, enabling research on privacy policies across different categories/market sectors. We review the statistics of this corpus and make it available for research. We also obtain valuable insights about privacy policies, e.g., which websites post them less often. Our corpus of web privacy policies is a valuable tool at the researchers' disposal to investigate privacy policies. For example, it facilitates comparison among different methods of privacy policy summarization by providing a benchmark, and can be used in unsupervised machine learning to summarize privacy policies. Razieh Nokhbeh Zaeem, K. Suzanne Barber |
CODASPY | 1 |
| 2021 | Comparing Privacy Policies of Government Agencies and Companies: A Study using Machine-learning-based Privacy Policy Analysis Tools
Razieh Nokhbeh Zaeem, K. Suzanne Barber |
ICAART (2) | 1 |
| 2020 | On Sentiment of Online Fake NewsabstractThe presence of disinformation and fake news on the Internet and especially social media has become a major concern. Prime examples of such fake news surged in the 2016 U.S. presidential election cycle and the COVID-19 pandemic. We quantify sentiment differences between true and fake news on social media using a diverse body of datasets from the literature that contains about 100K previously labeled true and fake news. We also experiment with a variety of sentiment analysis tools. We model the association between sentiment and veracity as conditional probability and also leverage statistical hypothesis testing to uncover the relationship between sentiment and veracity. With a significance level of 99.999%, we observe a statistically significant relationship between negative sentiment and fake news and between positive sentiment and true news. The degree of association, as measured by Goodman and Kruskal's gamma, ranges between. 037 to. 475. Finally, we make our data and code publicly available to support reproducibility. Our results assist in the development of automatic fake news detectors. Razieh Nokhbeh Zaeem, Chengjing Li, K. Suzanne Barber |
ASONAM | 1 |
| 2020 | PrivacyCheck v2: A Tool that Recaps Privacy Policies for YouabstractDespite the efforts to regulate privacy policies to protect user privacy, these policies remain lengthy and hard to comprehend. Powered by machine learning, our publicly available browser extension, PrivacyCheck v2, automatically summarizes any privacy policy by answering 20 questions based upon User Control and the General Data Protection Regulation. Furthermore, PrivacyCheck v2 incorporates a competitor analysis tool that highlights the top competitors with the best privacy policies in the same market sector. PrivacyCheck v2 enhances the users' understanding of privacy policies and empowers them to make informed decisions when it comes to selecting services with better privacy policies. Razieh Nokhbeh Zaeem, Safa Anya, Alex Issa, Jake Nimergood, Isabelle Rogers, Vinay Shah, Ayush Srivastava, K. Suzanne Barber |
CIKM | 1 |
| 2020 | An Evaluation Framework for Future Privacy Protection Systems: A Dynamic Identity Ecosystem Approach
David Liau, Razieh Nokhbeh Zaeem, K. Suzanne Barber |
ICAART (1) | 2 |
| 2020 | A Framework for Estimating Privacy Risk Scores of Mobile Apps
Kai Chih Chang, Razieh Nokhbeh Zaeem, K. Suzanne Barber |
ISC | 2 |
| 2019 | Evaluation Framework for Future Privacy Protection Systems: A Dynamic Identity Ecosystem ApproachabstractIn this paper, we leverage previous work in the Identity Ecosystem, a Bayesian network mathematical representation of a person's identity, to create a framework to evaluate identity protection systems. Information dynamic is considered and a protection game is formed given that the owner and the attacker both gain some level of control over the status of other PII within the dynamic Identity Ecosystem. We present a policy iteration algorithm to solve the optimal policy for the game and discuss its convergence. Finally, an evaluation and comparison of identity protection strategies is provided given that an optimal policy is used against different protection policies. This study is aimed to understand the evolutionary process of identity theft and provide a framework for evaluating different identity protection strategies and future privacy protection system. David Liau, Razieh Nokhbeh Zaeem, K. Suzanne Barber |
PST | 2 |
| 2019 | An Assessment of Blockchain Identity Solutions: Minimizing Risk and Liability of AuthenticationabstractPersonally Identifiable Information (PII) is often used to perform authentication and acts as a gateway to personal and organizational information. One weak link in the architecture of identity management services is sufficient to cause exposure and risk identity. Recently, we have witnessed a shift in identity management solutions with the growth of blockchain. Blockchain—the decentralized ledger system—provides a unique answer addressing security and privacy with its embedded immutability. In a blockchain-based identity solution, the user is given the control of his/her identity by storing personal information on his/her device and having the choice of identity verification document used later to create blockchain attestations. Yet, the blockchain technology alone is not enough to produce a better identity solution. The user cannot make informed decisions as to which identity verification document to choose if he/she is not presented with tangible guidelines. In the absence of scientifically created practical guidelines, these solutions and the choices they offer may become overwhelming and even defeat the purpose of providing a more secure identity solution. Rima Rana, Razieh Nokhbeh Zaeem, K. Suzanne Barber |
WI | 2 |
| 2019 | It Is an Equal Failing to Trust Everybody and to Trust Nobody: Stock Price Prediction Using Trust Filters and Enhanced User Sentiment on TwitterabstractSocial media are providing a huge amount of information, in scales never possible before. Sentiment analysis is a powerful tool that uses social media information to predict various target domains (e.g., the stock market). However, social media information may or may not come from trustworthy users. To utilize this information, a very first critical problem to solve is to filter credible and trustworthy information from contaminated data, advertisements, or scams. We investigate different aspects of a social media user to score his/her trustworthiness and credibility. Furthermore, we provide suggestions on how to improve trustworthiness on social media by analyzing the contribution of each trust score. We apply trust scores to filter the tweets related to the stock market as an example target domain. While social media sentiment analysis has been on the rise over the past decade, our trust filters enhance conventional sentiment analysis methods and provide more accurate prediction of the target domain, here, the stock market. We argue that while it is a failing to ignore the information social media provide, effectively trusting nobody, it is an equal failing to trust everybody on social media too: Our filters seek to identify whom to trust. Teng-Chieh Huang, Razieh Nokhbeh Zaeem, K. Suzanne Barber |
ACM Trans. Internet Techn. | 2 |
| 2018 | PrivacyCheck: Automatic Summarization of Privacy Policies Using Data MiningabstractPrior research shows that only a tiny percentage of users actually read the online privacy policies they implicitly agree to while using a website. Prior research also suggests that users ignore privacy policies because these policies are lengthy and, on average, require 2 years of college education to comprehend. We propose a novel technique that tackles this problem by automatically extracting summaries of online privacy policies. We use data mining models to analyze the text of privacy policies and answer 10 basic questions concerning the privacy and security of user data, what information is gathered from them, and how this information is used. In order to train the data mining models, we thoroughly study privacy policies of 400 companies (considering 10% of all listings on NYSE, Nasdaq, and AMEX stock markets) across industries. Our free Chrome browser extension, PrivacyCheck, utilizes the data mining models to summarize any HTML page that contains a privacy policy. PrivacyCheck stands out from currently available counterparts because it is readily applicable on any online privacy policy. Cross-validation results show that PrivacyCheck summaries are accurate 40% to 73% of the time. Over 400 independent Chrome users are currently using PrivacyCheck. Razieh Nokhbeh Zaeem, Rachel L. German, K. Suzanne Barber |
ACM Trans. Internet Techn. | 1 |
| 2017 | Automated Test Generation and Mutation Testing for AlloyabstractWe present two novel approaches for automated testing of models written in Alloy – a well-known declarative, first-order language that is supported by a fully automatic SAT-based analysis engine. The first approach introduces automated test generation for Alloy and is embodied by three techniques that create test suites in the traditional spirit of black-box, white-box, and mutation-based testing. The second approach introduces mutation testing for Alloy and defines how to create mutants of Alloy models, compute mutation testing results, and check for equivalent mutants using SAT. The two approaches build on the theoretical foundation defined previously by our AUnit framework, which introduced the idea of unit testing for Alloy in the spirit of unit testing for imperative languages. While test generation and mutation testing are heavily studied problems with many solutions in the context of imperative languages, the key novelty of our work is to introduce and address these problems for the declarative programming paradigm, specifically for the Alloy language. Experimental results using several Alloy subjects, including those with real faults, demonstrate the efficacy of our framework. Allison Sullivan, Razieh Nokhbeh Zaeem, Sarfraz Khurshid |
ICST | 3 |
| 2017 | Modeling and analysis of identity threat behaviors through text mining of identity theft stories
Razieh Nokhbeh Zaeem, Monisha Manoharan, K. Suzanne Barber |
Comput. Secur. | 1 |
| 2014 | Automated Generation of Oracles for Testing User-Interaction Features of Mobile AppsabstractAs the use of mobile devices becomes increasingly ubiquitous, the need for systematically testing applications (apps) that run on these devices grows more and more. However, testing mobile apps is particularly expensive and tedious, often requiring substantial manual effort. While researchers have made much progress in automated testing of mobile apps during recent years, a key problem that remains largely untracked is the classic oracle problem, i.e., to determine the correctness of test executions. This paper presents a novel approach to automatically generate test cases, that include test oracles, for mobile apps. The foundation for our approach is a comprehensive study that we conducted of real defects in mobile apps. Our key insight, from this study, is that there is a class of features that we term user-interaction features, which is implicated in a significant fraction of bugs and for which oracles can be constructed - in an application agnostic manner -- based on our common understanding of how apps behave. We present an extensible framework that supports such domain specific, yet application agnostic, test oracles, and allows generation of test sequences that leverage these oracles. Our tool embodies our approach for generating test cases that include oracles. Experimental results using 6 Android apps show the effectiveness of our tool in finding potentially serious bugs, while generating compact test suites for user-interaction features. Razieh Nokhbeh Zaeem, Mukul R. Prasad, Sarfraz Khurshid |
ICST | 1 |
| 2014 | Towards a test automation framework for alloyabstractWriting declarative models of software designs and analyzing them to detect defects is an effective methodology for developing more dependable software systems. However, writing such models correctly can be challenging for practitioners who may not be proficient in declarative programming, and their models themselves may be buggy. We introduce the foundations of a novel test automation framework, AUnit, which we envision for testing declarative models written in Alloy -- a first-order, relational language that is supported by its SAT-based analyzer. We take inspiration from the success of the family of xUnit frameworks that are used widely in practice for test automation, albeit for imperative or object-oriented programs. The key novelty of our work is to define a basis for unit testing for Alloy, specifically, to define the concepts of test case and coverage, and coverage criteria for declarative models. We reduce the problems of declarative test execution and coverage computation to evaluation without requiring SAT solving. Our vision is to blend how developers write unit tests in commonly used programming languages with how Alloy users formulate their models in Alloy, thereby facilitating the development and testing of Alloy models for both new Alloy users as well as experts. We illustrate our ideas using a small but complex Alloy model. While we focus on Alloy, our ideas generalize to other declarative languages (such as Z, B, ASM). Allison Sullivan, Razieh Nokhbeh Zaeem, Sarfraz Khurshid, Darko Marinov |
SPIN | 2 |
| 2013 | Repair Abstractions for More Efficient Data Structure Repair
Razieh Nokhbeh Zaeem, Muhammad Zubair Malik, Sarfraz Khurshid |
RV | 1 |
| 2012 | Improving the effectiveness of spectra-based fault localization using specificationsabstractFault localization i.e., locating faulty lines of code, is a key step in removing bugs and often requires substantial manual effort. Recent years have seen many automated localization techniques, specifically using the program’s passing and failing test runs, i.e., test spectra. However, the effectiveness of these approaches is sensitive to factors such as the type and number of faults, and the quality of the test-suite. This paper presents a novel technique that applies spectra-based localization in synergy with specification-based analysis to more accurately locate faults. Our insight is that unsatisfiability analysis of violated specifications, enabled by SAT technology, could be used to (1) compute unsatisfiable cores that contain likely faulty statements and (2) generate tests that help spectra-based localization. Our technique is iterative and driven by a feedback loop that enables more precise fault localization. SAT-TAR is a framework that embodies our technique for Java programs, including those with multiple faults. An experimental evaluation using a suite of widely-studied data structure programs, including the ANTLR and JTopas parser applications, shows that our technique localizes faults more accurately than state-of-the-art approaches. Divya Gopinath, Razieh Nokhbeh Zaeem, Sarfraz Khurshid |
ASE | 2 |
| 2012 | Test input generation using dynamic programmingabstractConstraint-based input generation is an effective technique for testing programs, such as compilers and web browsers, which have complex inputs. However, efficient generation of such inputs remains a challenging problem. We present a novel input generation technique that takes constraints written as recursive predicates in the underlying programming language and uses dynamic programming to solve the constraints efficiently. Our key insight is to leverage the recursive structure of desired inputs and partition the problem of generating an input into several sub-problems of generating smaller inputs that exhibit the same structure, and then to use dynamic programming -- a well-known problem solving methodology designed to exploit common sub-problems -- to combine them. A lazy initialization strategy and symbolic execution optimize our basic technique. Our technique provides not only bounded exhaustive input generation but also enables random input generation. We show the correctness of our technique. Furthermore, we present an experimental evaluation, which shows that our technique can provide over an order of magnitude performance improvement for input generation compared to Korat (an efficient solver for structural constraints) and Pex (a state-of-the-art tool for symbolic execution). Finally, we use our technique to effectively find bugs in production versions of Google Chrome and Apple Safari web browsers. Razieh Nokhbeh Zaeem, Sarfraz Khurshid |
SIGSOFT FSE | 1 |
| 2012 | History-Aware Data Structure Repair Using SAT
Razieh Nokhbeh Zaeem, Divya Gopinath, Sarfraz Khurshid, Kathryn S. McKinley |
TACAS | 1 |
| 2010 | Contract-Based Data Structure Repair Using Alloy
Razieh Nokhbeh Zaeem, Sarfraz Khurshid |
ECOOP | 1 |