Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Ran Wolff 0003

dblp:11/4475-3 · DBLP profile ↗
← Back
4ranked-venue papers
2as first author
0since 2021 · last 2016
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 2 · 1 first-authorArtificial intelligence and machine learning · 1Systems, architecture and hardware · 1 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1Human-computer interaction and ubiquitous computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Network and information security
1 paper
Privacy and data protection · 100%
Theoretical computer science
1 paper
Distributed computing theory · 100%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Distributed systems · 100%

Topics — the 3 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Privacy and data protection
anonymization
0.212016
Enforcing k-anonymity in Web Mail Auditing · WSDM 2016
Distributed computing theory › communication-efficient algorithms
communication-efficient distributed algorithms
0.212015
Distributed Convex Thresholding · PODC 2015
Privacy and data protection › data confidentiality › content privacy
email privacy
0.112016
Enforcing k-anonymity in Web Mail Auditing · WSDM 2016

Methods — techniques the papers use, named apart from their topics

convex thresholding · 0.4communication-efficient monitoring · 0.4message signature · 0.2equivalence class masking · 0.2
YearPublicationVenuePosition
2016 Enforcing k-anonymity in Web Mail Auditing
abstract
We study the problem of k-anonymization of mail messages in the realistic scenario of auditing mail traffic in a major commercial Web mail service. Mail auditing is necessary in various Web mail debugging and quality assurance activities, such as anti-spam or the qualitative evaluation of novel mail features. It is conducted by trained professionals, often referred to as "auditors", who are shown messages that could expose personally identifiable information. We address here the challenge of k-anonymizing such messages, focusing on machine generated mail messages that represent more than 90% of today's mail traffic. We introduce a novel message signature Mail-Hash, specifically tailored to identifying structurally-similar messages, which allows us to put such messages in a same equivalence class. We then define a process that generates, for each class, masked mail samples that can be shown to auditors, while guaranteeing the k-anonymity of users. The productivity of auditors is measured by the amount of non-hidden mail content they can see every day, while considering normal working conditions, which set a limit to the number of mail samples they can review. In addition, we consider k-anonymity over time since, by definition of k-anonymity, every new release places additional constraints on the assignment of samples. We describe in details the results we obtained over actual Yahoo mail traffic, and thus demonstrate that our methods are feasible at Web mail scale. Given the constantly growing concern of users over their email being scanned by others, we argue that it is critical to devise such algorithms that guarantee k-anonymity, and implement associated processes in order to restore the trust of mail users.
Dotan Di Castro, Liane Lewin-Eytan, Yoelle Maarek, Ran Wolff 0003, Eyal Zohar
WSDM4
2015 Plot Balalaika: Simple Chart Designs for Long-Tail Distributed Data
abstract
Current approaches to summarising large arrays of data for presentation and communication mostly comprise reporting means with, e.g., Bar-charts. These methods are well-suited for unimodal, ideally normally-or near-normally distributed data, but are misleading for long-tail distributions that comprise most of the Big Data. We propose a succinct visualisation format, parallel in simplicity to bar-charts, that is suitable for communicating the gist of long-tail distributions, and show its efficiency empirically.
Mark M. Shovman, Ran Wolff 0003
IV2
2015 Distributed Convex Thresholding
abstract
Over the last fifteen years, a large group of algorithms emerged which compute various predicates from distributed data with a focus on communication efficiency. These algorithms are often called "communication-efficient", "geometric-monitoring", or "local" algorithms. We jointly call them distributed convex thresholding algorithms, for reasons which will be explained in this work. Distributed convex thresholding algorithms have found their applications in domains in which bandwidth is a scarce resource, such as wireless sensor networks and peer-to-peer systems, or in scenarios in which data rapidly streams to the different processors but outcome of the predicate rarely changes. Common to all of these algorithms is the use of a data dependent criteria to determine when further messaging is required.
Ran Wolff 0003
PODC1
2006 Local L2-Thresholding Based Data Mining in Peer-to-Peer Systems
abstract
In a large network of computers, wireless sensors, or mobile devices, each of the components (hence, peers) has some data about the global status of the system. Many of the functions of the system, such as routing decisions, search strategies, data cleansing, and the assignment of mutual trust, depend on the global status. Therefore, it is essential that the system be able to detect, and react to, changes in its global status. Computing global predicates in such systems is usually very costly. Mainly because of their scale, and in some cases (e.g., sensor networks) also because of the high cost of communication. The cost further increases when the data changes rapidly (due to state changes, node failure, etc.) and computation has to follow these changes. In this paper we describe a two step approach for dealing with these costs. First, we describe a highly efficient local algorithm which detect when the L2 norm of the average data surpasses a threshold. Then, we use this algorithm as a feedback loop for the monitoring of complex predicates on the data - such as the data's k-means clustering. The efficiency of the L2 algorithm guarantees that so long as the clustering results represent the data (i.e., the data is stationary) few resources are required. When the data undergoes an epoch change - a change in the underlying distribution - and the model no longer represents it, the feedback loop indicates this and the model is rebuilt. Furthermore, the existence of a feedback loop allows using approximate and “best-effort” methods for constructing the model; if an ill-fit model is built the feedback loop would indicate so, and the model would be rebuilt.
Ran Wolff 0003, Kanishka Bhaduri, Hillol Kargupta
SDM1