VLDB 2026 Research / reviewers in the wild / expert
Ilana Segall
dblp:183/6384
· DBLP profile ↗
1ranked-venue papers
0as first author
0since 2021 · last 2020
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 1Applied, interdisciplinary, general and emerging computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Network and information security
1 paper |
Privacy and data protection · 100% | |
| Computer networks
1 paper |
Network measurement and analytics · 100% |
Topics — the 3 heaviest of 3, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Network measurement and analytics
web measurement |
0.4 | 1 | 2020 | The Representativeness of Automated Web Crawls as a Surrogate for Human Browsing · WWW 2020 |
Privacy and data protection
online tracking |
0.4 | 1 | 2020 | The Representativeness of Automated Web Crawls as a Surrogate for Human Browsing · WWW 2020 |
Privacy and data protection › web tracking
browser fingerprinting |
0.1 | 1 | 2020 | The Representativeness of Automated Web Crawls as a Surrogate for Human Browsing · WWW 2020 |
Methods — techniques the papers use, named apart from their topics
measurement study · 0.9comparative analysis · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2020 | The Representativeness of Automated Web Crawls as a Surrogate for Human BrowsingabstractLarge-scale Web crawls have emerged as the state of the art for studying characteristics of the Web. In particular, they are a core tool for online tracking research. Web crawling is an attractive approach to data collection, as crawls can be run at relatively low infrastructure cost and don’t require handling sensitive user data such as browsing histories. However, the biases introduced by using crawls as a proxy for human browsing data have not been well studied. Crawls may fail to capture the diversity of user environments, and the snapshot view of the Web presented by one-time crawls does not reflect its constantly evolving nature, which hinders reproducibility of crawl-based studies. In this paper, we quantify the repeatability and representativeness of Web crawls in terms of common tracking and fingerprinting metrics, considering both variation across crawls and divergence from human browser usage. We quantify baseline variation of simultaneous crawls, then isolate the effects of time, cloud IP address vs. residential, and operating system. This provides a foundation to assess the agreement between crawls visiting a standard list of high-traffic websites and actual browsing behaviour measured from an opt-in sample of over 50,000 users of the Firefox Web browser. Our analysis reveals differences between the treatment of stateless crawling infrastructure and generally stateful human browsing, showing, for example, that crawlers tend to experience higher rates of third-party activity than human browser users on loading pages from the same domains. David Zeber, Sarah Bird, Camila Oliveira, Walter Rudametkin, Ilana Segall, Fredrik Wollsén, Martin Lopatka |
WWW | 5 |