Hamed Haddadi 0001

dblp:33/5454-1 · DBLP profile ↗
← Back
80ranked-venue papers
0as first author
45since 2021 · last 2026
0000-0002-5895-8903ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 31 · 13 since 2021Security and privacy · 24 · 19 since 2021Artificial intelligence and machine learning · 10 · 8 since 2021Human-computer interaction and ubiquitous computing · 8Databases, data management, data science and information retrieval · 5 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 since 2021Systems, architecture and hardware · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 CAEC: Confidential, Attestable, and Efficient Inter-CVM Communication with Arm CCA
abstract
Confidential Virtual Machines (CVMs) are increasingly adopted to protect sensitive workloads from privileged adversaries such as the hypervisor. While they provide strong isolation guarantees, existing CVM architectures lack first-class mechanisms for inter-CVM data sharing due to their disjoint memory model, making inter-CVM data exchange a performance bottleneck in compartmentalized or collaborative multi-CVM systems. Under this model, a CVM's accessible memory is either shared with the hypervisor or protected from both the hypervisor and all other CVMs. This design simplifies reasoning about memory ownership; however, it fundamentally precludes plaintext data sharing between CVMs because all inter-CVM communication must pass through hypervisor-accessible memory, requiring costly encryption and decryption to preserve confidentiality and integrity. In this paper, we introduce CAEC, a system that enables protected memory sharing between CVMs. CAEC builds on Arm Confidential Compute Architecture (CCA) and extends its firmware to support Confidential Shared Memory (CSM), a memory region securely shared between multiple CVMs while remaining inaccessible to the hypervisor and all non-participating CVMs. CAEC's design is fully compatible with CCA hardware and introduces only a modest increase (6%) in CCA firmware code size. CAEC delivers substantial performance benefits across a range of workloads. For instance, inter-CVM communication over CAEC achieves up to 209x reduction in CPU cycles compared to encryption-based mechanisms over hypervisor-accessible shared memory. By combining high performance, strong isolation guarantees, and attestable sharing semantics, CAEC provides a practical and scalable foundation for the next generation of trusted multi-CVM services across both edge and cloud environments.
Sina Abdollahi, Amir Al Sadi, David Kotz, Marios Kogias, Hamed Haddadi 0001
EuroS&P5
2026 Short Paper: Towards Real-Time ECG and EMG Modeling on μNPUs
abstract
The miniaturisation of neural processing units (NPUs) and other low-power accelerators has enabled their integration into microcontroller-scale wearable hardware, supporting near-real-time, offline, and privacy-preserving inference. Yet physiological signal analysis has remained infeasible on such hardware; recent Transformer-based models show state-of-the-art performance but are prohibitively large for resource- and power-constrained hardware and incompatible with µNPUs due to their dynamic attention operations. We introduce PhysioLite, a lightweight, NPU-compatible model architecture and training framework for ECG/EMG signal analysis. Using learnable wavelet filter banks, CPU-offloaded positional encoding, and hardware-aware layer design, PhysioLite reaches performance comparable to state-of-the-art Transformer-based foundation models on ECG and EMG benchmarks, while being <10% of the size (∼ 370KB with 8-bit quantization). We also profile its component-wise latency and resource consumption on both the MAX78000 and HX6538 WE2 µNPUs, demonstrating its viability for signal analysis on constrained, battery-powered hardware. We release our model(s) and training framework at: https://github.com/j0shmillar/physiolite.
Josh Millar, Ashok Samraj Thangarajan, Soumyajit Chatterjee, Hamed Haddadi 0001
SenSys4
2026 Dynamic Probabilistic Noise Injection for Membership Inference Defense
abstract
Membership Inference Attacks (MIAs) expose privacy risks by determining whether a specific sample was part of a model’s training set. These threats are especially serious in sensitive domains such as healthcare and finance. Traditional mitigation techniques, such as static differential privacy, rely on injecting a fixed amount of noise during training or inference. However, this often leads to a detrimental trade-off: the noise may be insufficient to counter sophisticated attacks or, when increased, can substantially degrade model accuracy. To address this limitation, we propose DynaNoise, an adaptive inference-time defense that modulates injected noise based on per-query sensitivity. DynaNoise estimates risk using measures such as Shannon entropy and scales the noise variance accordingly, followed by a smoothing step that re-normalizes the perturbed outputs to preserve predictive utility. We further introduce MIDPUT (Membership Inference Defense Privacy-Utility Trade-off), a scalar metric that captures both privacy gains and accuracy retention. Our evaluation on several benchmark datasets demonstrates that DynaNoise substantially lowers attack success rates while maintaining competitive accuracy, achieving strong overall MIDPUT scores compared to state-of-the-art defenses.
Javad Forough, Hamed Haddadi 0001
Proc. Priv. Enhancing Technol.2
2025 Nebula: Efficient, Private and Accurate Histogram Estimation
abstract
International audience
Ali Shahin Shamsabadi, Peter Snyder, Ralph Giles, Aurélien Bellet, Hamed Haddadi 0001
CCS5
2025 Local Frames: Exploiting Inherited Origins to Bypass Content Blockers
abstract
We present a study of how local frames (i.e., iframes loading content like ''about:blank'') are mishandled by a wide range of popular Web security and privacy tools. As a result, users of these tools remain vulnerable to the very attack techniques against which they seek to protect themselves, including browser fingerprinting, cookie-based tracking, and data exfiltration. The tools we study are vulnerable in different ways, but all share a root cause: legacy Web functionality interacts with browser privacy boundaries in unexpected ways, leading to systemic vulnerabilities in tools developed, maintained, and recommended by privacy experts and activists.
Alisha Ukani, Hamed Haddadi 0001, Alex C. Snoeren, Peter Snyder
CCS2
2025 Context-Aware Membership Inference Attacks against Pre-trained Large Language Models
abstract
Membership Inference Attacks (MIAs) on pretrained Large Language Models (LLMs) aim at determining if a data point was part of the model's training set.Prior MIAs that are built for classification models fail at LLMs, due to ignoring the generative nature of LLMs across token sequences.In this paper, we present a novel attack on pre-trained LLMs that adapts MIA statistical tests to the perplexity dynamics of subsequences within a data point.Our method significantly outperforms prior approaches, revealing context-dependent memorization patterns in pre-trained LLMs.
Hongyan Chang, Ali Shahin Shamsabadi, Kleomenis Katevas, Hamed Haddadi 0001, Reza Shokri
EMNLP4
2025 Membership and Memorization in LLM Knowledge Distillation
abstract
Recent advances in Knowledge Distillation (KD) aim to mitigate the high computational demands of Large Language Models (LLMs) by transferring knowledge from a large "teacher" to a smaller "student" model.However, students may inherit the teacher's privacy when the teacher is trained on private data.In this work, we systematically characterize and investigate membership and memorization privacy risks inherent in six LLM KD techniques.Using instruction-tuning settings that span seven NLP tasks, together with three teacher model families (GPT-2, LLAMA-2, and OPT), and various size student models, we demonstrate that all existing LLM KD approaches carry membership and memorization privacy risks from the teacher to its students.However, the extent of privacy risks varies across different KD techniques.We systematically analyse how key LLM KD components (KD objective functions, student training data and NLP tasks) impact such privacy risks.We also demonstrate a significant disagreement between memorization and membership privacy risks of LLM KD techniques.Finally, we characterize per-block privacy risk and demonstrate that the privacy risk varies across different blocks by a large margin.
Ziqi Zhang 0017, Ali Shahin Shamsabadi, Hanxiao Lu, Yifeng Cai, Hamed Haddadi 0001
EMNLP5
2025 Deep Unlearn: Benchmarking Machine Unlearning for Image Classification
abstract
Machine unlearning (MU) aims to remove the influence of particular data points from the learnable parameters of a trained machine learning model. This is a crucial capability in light of data privacy requirements, trustworthiness, and safety in deployed models. MU is particularly challenging for deep neural networks (DNNs), such as convolutional nets or vision transformers, as such DNNs tend to memorize a notable portion of their training dataset. Nevertheless, the community lacks a rigorous and multifaceted study that looks into the success of MU methods for DNNs. In this paper, we investigate 18 state-of-the-art MU methods across various benchmark datasets and models, with each evaluation conducted over 10 different initializations, a comprehensive evaluation involving MU over 100K models. We show that, with the proper hyperparameters, Masked Small Gradients (MSG) and Convolution Transpose (CT), consistently perform better in terms of model accuracy and run-time efficiency across different models, datasets, and initializations, assessed by population-based membership inference attacks (MIA) and per-sample unlearning likelihood ratio attacks (U-LiRA). Furthermore, our benchmark highlights the fact that comparing a MU method only with commonly used baselines, such as Gradient Ascent (GA) or Successive Random Relabeling (SRL), is inadequate, and we need better baselines like Negative Gradient Plus (NG+) with proper hyperparameter selection.
Xavier F. Cadet, Anastasia Borovykh, Mohammad Malekzadeh, Sara Ahmadi-Abhari, Hamed Haddadi 0001
EuroS&P5
2025 ACCESS-FL: Agile Communication and Computation for Efficient Secure Aggregation in Stable Networks for FLaaS
abstract
Federated Learning (FL) enables privacy-preserving machine learning by allowing clients to collaboratively train models without sharing raw data. Federated Learning as a Service (FLaaS) extends this approach to cloud infrastructures. However, conventional secure aggregation protocols, such as Google's SecAgg and SecAgg+, introduce high computation and communication overheads, particularly in large-scale FLaaS deployments where client dropout rates are limited. To address these challenges, we propose ACCESS-FL, a lightweight, secure aggregation method designed for honest-but-curious FLaaS scenarios with stable network conditions. ACCESS-FL eliminates double masking, Shamir's Secret Sharing, and excessive encryption/decryption by creating shared secrets only between two peers per client, which reduces computation and communication complexity to constant$O(1)$and makes the algorithm independent of network size and comparable to standard FL. ACCESS-FL preserves privacy against inversion attacks and maintains model accuracy equivalent to the FL, SecAgg, and SecAgg+ protocols, proving that reducing overhead does not compromise learning performance and achieves communication and computation costs comparable to standard FL. Experimental evaluations on benchmark datasets (MNIST, FMNIST, and CIFAR-10) demonstrate lower overhead, making ACCESS-FL practical for service-based stable FLaaS applications such as healthcare analytics.
Niousha Nazemi, Omid Tavallaie, Shuaijun Chen, Anna Maria Mandalari, Kanchana Thilakarathna, Ralph Holz, Hamed Haddadi 0001, Albert Y. Zomaya
ICWS7
2025 Benchmarking Ultra-Low-Power μNPUs
abstract
Efficient on-device neural network (NN) inference offers predictable latency, improved privacy and reliability, and lower operating costs for vendors than cloud-based inference. This has sparked recent development of microcontroller-scale NN accelerators, also known as neural processing units (μNPUs), designed specifically for ultra-low-power applications.
Josh Millar, Yushan Huang, Sarab S. Sethi, Hamed Haddadi 0001, Anil Madhavapeddy
MobiCom4
2025 DiStefano: Decentralized Infrastructure for Sharing Trusted Encrypted Facts and Nothing More
Sofía Celi, Alex Davidson, Hamed Haddadi 0001, Gonçalo Pestana, Joe Rowell
NDSS3
2025 Secure and Confidential Certificates of Online Fairness
abstract
The "black-box service model" enables ML service providers to serve clients while keeping their intellectual property and client data confidential. Confidentiality is critical for delivering ML services legally and responsibly, but makes it difficult for outside parties to verify important model properties such as fairness. Existing methods that assess model fairness confidentially lack either (i) *reliability* because they certify fairness with respect to a static set of data, and therefore fail to guarantee fairness in the presence of distribution shift or service provider malfeasance; and/or (ii) *scalability* due to the computational overhead of confidentiality-preserving cryptographic primitives. We address these problems by introducing *online fairness certificates*, which verify that a model is fair with respect to data received by the service provider *online* during deployment. We then present OATH, a deployably efficient and scalable zero-knowledge proof protocol for confidential online group fairness certification. OATH exploits statistical properties of group fairness via a "cut-and-choose" style protocol, enabling scalability improvements over baselines.
Olive Franzese-McLaughlin, Ali Shahin Shamsabadi, Carter Luck, Hamed Haddadi 0001
NeurIPS4
2025 Robust Hallucination Detection in LLMs via Adaptive Token Selection
abstract
Hallucinations in large language models (LLMs) pose significant safety concerns that impede their broader deployment. Recent research in hallucination detection has demonstrated that LLMs' internal representations contain truthfulness hints, which can be harnessed for detector training. However, the performance of these detectors is heavily dependent on the internal representations of predetermined tokens, fluctuating considerably when working on free-form generations with varying lengths and sparse distributions of hallucinated entities. To address this, we propose HaMI, a novel approach that enables robust detection of hallucinations through adaptive selection and learning of critical tokens that are most indicative of hallucinations. We achieve this robustness by an innovative formulation of the Hallucination detection task as Multiple Instance (HaMI) learning over token-level representations within a sequence, thereby facilitating a joint optimisation of token selection and hallucination detection on generation sequences of diverse forms. Comprehensive experimental results on four hallucination benchmarks show that HaMI significantly outperforms existing state-of-the-art approaches.
Mengjia Niu, Hamed Haddadi 0001, Guansong Pang
NeurIPS2
2025 Pin the Tail on the Model: Blindfolded Repair of User-Flagged Failures in Text-to-Image Services
abstract
Diffusion models are increasingly deployed in real-world text-to-image services. These models, however, encode implicit assumptions about the world based on web-scraped image-caption pairs used during training. Over time, such assumptions may become outdated, incorrect, or socially biased--leading to failures where the generated images misalign with users' expectations or evolving societal norms. Identifying and fixing such failures is challenging and, thus, a valuable asset for service providers, as failures often emerge post-deployment and demand specialized expertise and resources to resolve them. In this work, we introduce $\textit{SURE}$, the first end‑to‑end framework that $\textbf{S}$ec$\textbf{U}$rely $\textbf{RE}$pairs failures flagged by users of diffusion-based services. $\textit{SURE}$ enables the service provider to securely collaborate with an external third-party specialized in model repairing (i.e., Model Repair Institute) without compromising the confidentiality of user feedback, the service provider’s proprietary model, or the Model Repair Institute’s proprietary repairing knowledge. To achieve the best possible efficiency, we propose a co-design of a model editing algorithm with a customized two-party cryptographic protocol. Our experiments show that $\textit{SURE}$ is highly practical: $\textit{SURE}$ securely and effectively repairs all 32 layers of {Stable Diffusion v1.4} in under 17 seconds (four orders of magnitude more efficient than a general baseline). Our results demonstrate that practical, secure model repair is attainable for large-scale, modern diffusion services.
Gefei Tan, Ali Shahin Shamsabadi, Ellen Kolesnikova, Hamed Haddadi 0001, Xiao Wang 0012
NeurIPS4
2025 Web Execution Bundles: Reproducible, Accurate, and Archivable Web Measurements
Florian Hantke, Peter Snyder, Hamed Haddadi 0001, Ben Stock
USENIX Security Symposium3
2025 Measuring the Accuracy and Effectiveness of PII Removal Services
abstract
This paper presents the first large-scale empirical study of commercial personally identifiable information (PII) removal systems --- commercial services that claim to improve privacy by automating the removal of PII from data broker's databases. Popular examples of such services include DeleteMe, Mozilla Monitor, Incogni, among many others. The claims these services make may be very appealing to privacy-conscious Web users, but how effective these services actually are at improving privacy has not been investigated. This work aims to improve our understanding of commercial PII removal services in multiple ways. First, we conduct a user study where participants purchase subscriptions from four popular PII removal services, and report (i) what PII the service find, (ii) from which data brokers, (iii) whether the service is able to have the information removed, and (iv) whether the identified information actually is PII describing the participant. And second, by comparing the claims and promises the services makes (e.g. which and how many data brokers each service claims to cover). We find that these services have significant accuracy and coverage issues that limit the usefulness of these services as a privacy-enhancing technology. For example, we find that the measured services are unable to remove the majority of the identified PII records from data broker's (48.2% of the successfully removed found records) and that most records identified by these services are not PII about the user (study participants found that only 41.1% of records identified by these services were PII about themselves).
Jiahui He 0001, Peter Snyder, Hamed Haddadi 0001, Fabián E. Bustamante, Gareth Tyson
Proc. Priv. Enhancing Technol.3
2025 TeleSparse: Practical Privacy-Preserving Verification of Deep Neural Networks
abstract
Verification of the integrity of deep learning inference is crucial for understanding whether a model is being applied correctly. However, such verification typically requires access to model weights and (potentially sensitive or private) training data. So-called Zero-knowledge Succinct Non-Interactive Arguments of Knowledge (ZK-SNARKs) would appear to provide the capability to verify model inference without access to such sensitive data. However, applying ZK-SNARKs to modern neural networks, such as transformers and large vision models, introduces significant computational overhead. We present TeleSparse, a ZK-friendly post-processing mechanisms to produce practical solutions to this problem. TeleSparse tackles two fundamental challenges inherent in applying ZK-SNARKs to modern neural networks: (1) Reducing circuit constraints: Over-parameterized models result in numerous constraints for ZK-SNARK verification, driving up memory and proof generation costs. We address this by applying sparsification to neural network models, enhancing proof efficiency without compromising accuracy or security. (2) Minimizing the size of lookup tables required for non-linear functions, by optimizing activation ranges through neural teleportation, a novel adaptation for narrowing activation functions’ range. TeleSparse reduces prover memory usage by 67% and proof generation time by 46% on the same model, with an accuracy trade-off of approximately 1%. We implement our framework using the Halo2 proving system and demonstrate its effectiveness across multiple architectures (Vision-transformer, ResNet, MobileNet) and datasets (ImageNet,CIFAR-10,CIFAR-100). This work opens new directions for ZK-friendly model design, moving toward scalable, resource-efficient verifiable deep learning.
Mohammad Maheri, Hamed Haddadi 0001, Alex Davidson
Proc. Priv. Enhancing Technol.2
2024 Unbundle-Rewrite-Rebundle: Runtime Detection and Rewriting of Privacy-Harming Code in JavaScript Bundles
abstract
This work presents Unbundle-Rewrite-Rebundle (URR), a system for detecting privacy-harming portions of bundled JavaScript code and rewriting that code at runtime to remove the privacy-harming behavior without breaking the surrounding code or overall application. URR is a novel solution to the problem of JavaScript bundles, where websites pre-compile multiple code units into a single file, making it impossible for content filters and ad-blockers to differentiate between desired and unwanted resources. Where traditional content filtering tools rely on URLs, URR analyzes the code at the AST level, and replaces harmful AST sub-trees with privacy-and-functionality maintaining alternatives.
Mir Masood Ali, Peter Snyder, Chris Kanich, Hamed Haddadi 0001
CCS4
2024 Confidential-DPproof: Confidential Proof of Differentially Private Training
abstract
Post hoc privacy auditing techniques can be used to test the privacy guarantees of a model, but come with several limitations: (i) they can only establish lower bounds on the privacy loss, (ii) the intermediate model updates and some data must be shared with the auditor to get a better approximation of the privacy loss, and (iii) the auditor typically faces a steep computational cost to run a large number of attacks. In this paper, we propose to proactively generate a cryptographic certificate of privacy during training to forego such auditing limitations. We introduce Confidential-DPproof , a framework for Confidential Proof of Differentially Private Training, which enhances training with a certificate of the $(\varepsilon,\delta)$-DP guarantee achieved. To obtain this certificate without revealing information about the training data or model, we design a customized zero-knowledge proof protocol tailored to the requirements introduced by differentially private training, including random noise addition and privacy amplification by subsampling. In experiments on CIFAR-10, Confidential-DPproof trains a model achieving state-of-the-art $91$% test accuracy with a certified privacy guarantee of $(\varepsilon=0.55,\delta=10^{-5})$-DP in approximately 100 hours.
Ali Shahin Shamsabadi, Gefei Tan, Tudor Cebere, Aurélien Bellet, Hamed Haddadi 0001, Nicolas Papernot, Xiao Wang 0012, Adrian Weller
ICLR5
2024 Low-Energy On-Device Personalization for MCUs
abstract
Microcontroller Units (MCUs) are ideal platforms for edge applications due to their low cost and energy consumption, and are widely used in various applications, including personalized machine learning tasks, where customized models can enhance the task adaptation. However, existing approaches for local on-device personalization mostly support simple ML architectures or require complex local pre-training/training, leading to high energy consumption and negating the low-energy advantage of MCUs. In this paper, we introduce MicroT, an efficient and lowenergy MCU personalization approach. MicroT includes a robust, general, but tiny feature extractor, developed through self-supervised knowledge distillation, which trains a task-specific head to enable independent on-device personalization with minimal energy and computational requirements. MicroT implements an MCU-optimized early-exit inference mechanism called stage-decision to further reduce energy costs. This mechanism allows for user-configurable exit criteria (stage-decision ratio) to adaptively balance energy cost with model performance. We evaluated MicroT using two models, three datasets, and two MCD boards. MicroT outperforms traditional transfer learning (TTL) and two SOTA approaches by 2.12 – 11.60% across two models and three datasets. Targeting widely used energy-aware edge devices, MicroT's on-device training requires no additional complex operations, halving the energy cost compared to SOTA approaches by up to 2.28× while keeping SRAM usage below 1MB. During local inference, MicroT reduces energy cost by 14.17% compared to TTL across two boards and two datasets, highlighting its suitability for long-term use on energy-aware resource-constrained MCUs.
Yushan Huang, Ranya Aloufi, Xavier F. Cadet, Payam M. Barnaghi, Hamed Haddadi 0001
SEC6
2024 A First Look at Related Website Sets
abstract
We present the first measurement of the user-effect and privacy impact of "Related Website Sets," a recent proposal to reduce browser privacy protections between two sites if those sites are related to each other. An assumption (both explicitly and implicitly) underpinning the Related Website Sets proposal is that users can accurately determine if two sites are related via the same entity. In this work, we probe this assumption via measurements and a user study of 30 participants, to assess the ability of Web users to determine if two sites are (according to the Related Website Sets feature) related to each other. We find that this is largely not the case. Our findings indicate that 42 (36.8%) of the user determinations in our study are incorrect in privacy-harming ways, where users think that sites are not related, but would be treated as related (and so due less privacy protections) by the Related Website Sets feature. Additionally, 22 (73.3%) of participants made at least one incorrect evaluation during the study. We also characterise the Related Website Sets list, its composition over time, and its governance.
Stephen McQuistin, Peter Snyder, Hamed Haddadi 0001, Gareth Tyson
IMC3
2024 MELTing Point: Mobile Evaluation of Language Transformers
abstract
Transformers have recently revolutionized the machine learning (ML) landscape, gradually making their way into everyday tasks and equipping our computers with "sparks of intelligence". However, their runtime requirements have prevented them from being broadly deployed on mobile. As personal devices become increasingly powerful at the consumer edge and prompt privacy becomes an ever more pressing issue, we explore the current state of mobile execution of Large Language Models (LLMs). To achieve this, we have created our own automation infrastructure, MELT, which supports the headless execution and benchmarking of LLMs on device, supporting different models, devices and frameworks, including Android, iOS and Nvidia Jetson devices. We evaluate popular instruction fine-tuned LLMs and leverage different frameworks to measure their end-to-end and granular performance, tracing their memory and energy requirements along the way.
Stefanos Laskaridis, Kleomenis Katevas, Lorenzo Minto, Hamed Haddadi 0001
MobiCom4
2024 SunBlock: Cloudless Protection for IoT Systems
Vadim Safronov, Anna Maria Mandalari, Daniel J. Dubois, David R. Choffnes, Hamed Haddadi 0001
PAM (2)5
2024 Poster: Towards Low-Power Comprehensive Biodiversity Monitoring
abstract
The Kunming-Montreal Global Biodiversity Framework sets ambitious targets for 2023, including halting human-induced species extinction. Achieving these requires comprehensive data on global biodiversity patterns, which can only be gathered through in-situ distributed sensor networks. However, these multi-device networks are constrained by battery lifetimes, must gather rich data from power-hungry sensors, and yet need to be deployed in remote environments for long periods. This note introduces a prototype multi-sensor device, and outlines how embedded scheduling could be used for extending sensor lifetime and resource-efficiency.
Josh Millar, Sarab S. Sethi, Hamed Haddadi 0001, Anil Madhavapeddy
SenSys3
2024 Analyzing entropy features in time-series data for pattern recognition in neurological conditions
abstract
In the field of medical diagnosis and patient monitoring, effective pattern recognition in neurological time-series data is essential. Traditional methods predominantly based on statistical or probabilistic learning and inference often struggle with multivariate, multi-source, state-varying, and noisy data while also posing privacy risks due to excessive information collection and modeling. Furthermore, these methods often overlook critical statistical information, such as the distribution of data points and inherent uncertainties. To address these challenges, we introduce an information theory-based pipeline that leverages specialized features to identify patterns in neurological time-series data while minimizing privacy risks. We incorporate various entropy methods based on the characteristics of different scenarios and entropy. For stochastic state transition applications, we incorporate Shannon's entropy, entropy rates, entropy production, and the von Neumann entropy of Markov chains. When state modeling is impractical, we select and employ approximate entropy, increment entropy, dispersion entropy, phase entropy, and slope entropy. The pipeline's effectiveness and scalability are demonstrated through pattern analysis in a dementia care dataset and also an epileptic and a myocardial infarction dataset. The results indicate that our information theory-based pipeline can achieve average performance improvements across various models on the recall rate, F1 score, and accuracy by up to 13.08 percentage points, while enhancing inference efficiency by reducing the number of model parameters by an average of 3.10 times. Thus, our approach opens a promising avenue for improved, efficient, and critical statistical information-considered pattern recognition in medical time-series data.
Yushan Huang, Alexander Capstick, Francesca Palermo, Hamed Haddadi 0001, Payam M. Barnaghi
Artif. Intell. Medicine5
2023 GT-TSCH: Game-Theoretic Distributed TSCH Scheduler for Low-Power IoT Networks
abstract
Time-Slotted Channel Hopping (TSCH) is a synchronous medium access mode of the IEEE 802.15.4e standard designed for providing low-latency and highly-reliable end-to-end communication. TSCH constructs a communication schedule by combining frequency channel hopping with Time Division Multiple Access (TDMA). In recent years, IETF designed several standards to define general mechanisms for the implementation of TSCH. However, the problem of updating the TSCH schedule according to the changes of the wireless link quality and node's traffic load left unresolved. In this paper, we use non-cooperative game theory to propose GT-TSCH, a distributed TSCH scheduler designed for low-power IoT applications. By considering selfish behavior of nodes in packet forwarding, GT-TSCH updates the TSCH schedule in a distributed approach with low control overhead by monitoring the queue length, the place of the node in the Directed Acyclic Graph (DAG) topology, the quality of the wireless link, and the data packet generation rate. We prove the existence and uniqueness of Nash equilibrium in our game model and we find the optimal number of TSCH Tx timeslots to update the TSCH slotframe. To examine the performance of our contribution, we implement GT-TSCH on Zolertia Firefly IoT motes and the Contiki-NG Operating System (OS). The evaluation results reveal that GT-TSCH improves performance in terms of throughput and end-to-end delay compared to the state-of-the-art method.
Omid Tavallaie, Seid Miad Zandavi, Hamed Haddadi 0001, Albert Y. Zomaya
ICDCS3
2023 Understanding the Privacy Risks of Popular Search Engine Advertising Systems
abstract
We present the first extensive measurement of the privacy properties of the advertising systems used by privacy-focused search engines. We propose an automated methodology to study the impact of clicking on search ads on three popularprivate search engines which have advertising-based business models: StartPage, Qwant, and DuckDuckGo, and we compare them to two dominant data-harvesting ones: Google and Bing. We investigate the possibility of third parties tracking users when clicking on ads by analyzing first-party storage, redirection domain paths, and requests sent before, when, and after the clicks.
Salim Chouaki, Oana Goga, Hamed Haddadi 0001, Peter Snyder
IMC3
2023 A First Look at the Privacy Harms of the Public Suffix List
abstract
The public suffix list is a community-maintained list of rules that can be applied to domain names to determine how they should be grouped into logical organizations or companies. We present the first large-scale measurement study of how the public suffix list is used by open-source software on the Web and the privacy harm resulting from projects using outdated versions of the list. We measure how often developers include out-of-date versions of the public suffix list in their projects, how old included lists are, and estimate the real-world privacy harm with a model based on a large-scale crawl of the Web. We find that incorrect use of the public suffix list is common in open-source software, and that at least 43 open-source projects use hard-coded, outdated versions of the public suffix list. These include popular, security-focused projects, such as password managers and digital forensics tools. We also estimate that, because of these out-of-date lists, these projects make incorrect privacy decisions for 1313 effective top-level domains (eTLDs), affecting 50,750 domains, by extrapolating from data gathered by the HTTP Archive project.
Stephen McQuistin, Peter Snyder, Colin Perkins, Hamed Haddadi 0001, Gareth Tyson
IMC4
2023 PRISM: Privacy Preserving Healthcare Internet of Things Security Management
abstract
Consumer healthcare Internet of Things (IoT) devices are gaining popularity in our homes and hospitals. These devices provide continuous monitoring at a low cost and can be used to augment high-precision medical equipment. However, major challenges remain in applying pre-trained global models for anomaly detection on smart health monitoring, for a diverse set of individuals that they provide care for. In this paper, we propose PRISM, an edge-based system for experimenting with in-home smart healthcare devices. We develop a rigorous methodology that relies on automated IoT experimentation. We use a rich real-world dataset from in-home patient monitoring from 44 households of People Living With Dementia (PLWD) over two years. Our results indicate that anomalies can be identified with accuracy up to 99% and mean training times as low as 0.88 seconds. While all models achieve high accuracy when trained on the same patient, their accuracy degrades when evaluated on different patients.
Savvas Hadjixenophontos, Anna Maria Mandalari, Hamed Haddadi 0001
ISCC4
2023 Effective Abnormal Activity Detection on Multivariate Time Series Healthcare Data
abstract
Multivariate time series (MTS) data collected from multiple sensors provide the potential for accurate abnormal activity detection in smart healthcare scenarios. However, anomalies exhibit diverse patterns and become unnoticeable in MTS data. Consequently, achieving accurate anomaly detection is challenging since we have to capture both temporal dependencies of time series and inter-relationships among variables. To address this problem, we propose a Residual-based Anomaly Detection approach, Rs-AD, for effective representation learning and abnormal activity detection. We evaluate our scheme on a real-world gait dataset and the experimental results demonstrate an F1 score of 0.839.
Mengjia Niu, Hamed Haddadi 0001
MobiCom3
2023 Poster: Towards Battery-Free Machine Learning Inference and Model Personalization on MCUs
abstract
Machine learning (ML) is moving towards edge devices. However, ML models with high computational demands and energy consumption pose challenges for ML inference in resource-constrained environments, such as the deep sea. To address these challenges, we propose a battery-free ML inference and model personalization pipeline for microcontroller units (MCUs). As an example, we performed fish image recognition in the ocean. We evaluated and compared the accuracy, runtime, power, and energy consumption of the model before and after optimization. The results demonstrate that, our pipeline can achieve 97.78% accuracy with 483.82 KB Flash, 70.32 KB RAM, 118 ms runtime, 4.83 mW power, and 0.57 mJ energy consumption on MCUs, reducing by 64.17%, 12.31%, 52.42%, 63.74%, and 82.67%, compared to the baseline. The results indicate the feasibility of battery-free ML inference on MCUs.
Yushan Huang, Hamed Haddadi 0001
MobiSys2
2023 Protected or Porous: A Comparative Analysis of Threat Detection Capability of IoT Safeguards
abstract
Consumer Internet of Things (IoT) devices are increasingly common, from smart speakers to security cameras, in homes. Along with their benefits come potential privacy and security threats. To limit these threats a number of commercial services have become available (IoT safeguards). The safeguards claim to provide protection against IoT privacy risks and security threats. However, the effectiveness and the associated privacy risks of these safeguards remains a key open question. In this paper, we investigate the threat detection capabilities of IoT safeguards for the first time. We develop and release an approach for automated safeguards experimentation to reveal their response to common security threats and privacy risks. We perform thousands of automated experiments using popular commercial IoT safeguards when deployed in a large IoT testbed. Our results indicate not only that these devices may be ineffective in preventing risks, but also their cloud interactions and data collection operations may introduce privacy risks for the households that adopt them.
Anna Maria Mandalari, Hamed Haddadi 0001, Daniel J. Dubois, David R. Choffnes
SP2
2023 Pool-Party: Exploiting Browser Resource Pools for Web Tracking
Peter Snyder, Soroush Karami, Arthur Edelstein, Benjamin Livshits, Hamed Haddadi 0001
USENIX Security Symposium5
2023 Federated and distributed learning applications for electronic health records and structured medical data: a scoping review
abstract
OBJECTIVES: Federated learning (FL) has gained popularity in clinical research in recent years to facilitate privacy-preserving collaboration. Structured data, one of the most prevalent forms of clinical data, has experienced significant growth in volume concurrently, notably with the widespread adoption of electronic health records in clinical practice. This review examines FL applications on structured medical data, identifies contemporary limitations, and discusses potential innovations. MATERIALS AND METHODS: We searched 5 databases, SCOPUS, MEDLINE, Web of Science, Embase, and CINAHL, to identify articles that applied FL to structured medical data and reported results following the PRISMA guidelines. Each selected publication was evaluated from 3 primary perspectives, including data quality, modeling strategies, and FL frameworks. RESULTS: Out of the 1193 papers screened, 34 met the inclusion criteria, with each article consisting of one or more studies that used FL to handle structured clinical/medical data. Of these, 24 utilized data acquired from electronic health records, with clinical predictions and association studies being the most common clinical research tasks that FL was applied to. Only one article exclusively explored the vertical FL setting, while the remaining 33 explored the horizontal FL setting, with only 14 discussing comparisons between single-site (local) and FL (global) analysis. CONCLUSIONS: The existing FL applications on structured medical data lack sufficient evaluations of clinically meaningful benefits, particularly when compared to single-site analyses. Therefore, it is crucial for future FL applications to prioritize clinical motivations and develop designs and methodologies that can effectively support and aid clinical practice and research.
Siqi Li 0004, Pinyan Liu, Gustavo G. Nascimento, Fabio Renato Manzolli Leite, Bibhas Chakraborty, Chuan Hong, Yilin Ning, Feng Xie 0004, Zhen Ling Teo, Daniel S. W. Ting, Hamed Haddadi 0001, Marcus Eng Hock Ong, Marco Aurélio Peres, Nan Liu 0003
J. Am. Medical Informatics Assoc.12
2023 Privacy-Aware Adversarial Network in Human Mobility Prediction
abstract
As mobile devices and location-based services are increasingly developed in different smart city scenarios and applications, many unexpected privacy leakages have arisen due to geolocated data collection and sharing. User re-identification and other sensitive inferences are major privacy threats when geolocated data are shared with cloud-assisted applications. Significantly, four spatio-temporal points are enough to uniquely identify 95% of the individuals, which exacerbates personal information leakages. To tackle malicious purposes such as user re-identification, we propose an LSTM-based adversarial mechanism with representation learning to attain a privacy-preserving feature representation of the original geolocated data (i.e., mobility data) for a sharing purpose. These representations aim to maximally reduce the chance of user re-identification and full data reconstruction with a minimal utility budget (i.e., loss). We train the mechanism by quantifying privacy-utility trade-off of mobility datasets in terms of trajectory reconstruction risk, user re-identification risk, and mobility predictability. We report an exploratory analysis that enables the user to assess this trade-off with a specific loss function and its weight parameters. The extensive comparison results on four representative mobility datasets demonstrate the superiority of our proposed architecture in mobility privacy protection and the efficiency of the proposed privacy-preserving features extractor. We show that the privacy of mobility traces attains decent protection at the cost of marginal mobility utility. Our results also show that by exploring the Pareto optimal setting, we can simultaneously increase both privacy (45%) and utility (32%).
Yuting Zhan, Hamed Haddadi 0001, Afra J. Mashhadi
Proc. Priv. Enhancing Technol.2
2023 Paralinguistic Privacy Protection at the Edge
abstract
Voice user interfaces and digital assistants are rapidly entering our lives and becoming singular touch points spanning our devices. These always-on services capture and transmit our audio data to powerful cloud services for further processing and subsequent actions. Our voices and raw audio signals collected through these devices contain a host of sensitive paralinguistic information that is transmitted to service providers regardless of deliberate or false triggers. As our emotional patterns and sensitive attributes like our identity, gender, and well-being are easily inferred using deep acoustic models, we encounter a new generation of privacy risks by using these services. One approach to mitigate the risk of paralinguistic-based privacy breaches is to exploit a combination of cloud-based processing with privacy-preserving, on-device paralinguistic information learning and filtering before transmitting voice data. In this article we introduce EDGY , a configurable, lightweight, disentangled representation learning framework that transforms and filters high-dimensional voice data to identify and contain sensitive attributes at the edge prior to offloading to the cloud. We evaluate EDGY’s on-device performance and explore optimization techniques, including model quantization and knowledge distillation, to enable private, accurate, and efficient representation learning on resource-constrained devices. Our results show that EDGY runs in tens of milliseconds with 0.2% relative improvement in “zero-shot” ABX score or minimal performance penalties of approximately 5.95% word error rate (WER) in learning linguistic representations from raw voice signals, using a CPU and a single-core ARM processor without specialized hardware.
Ranya Aloufi, Hamed Haddadi 0001, David Boyle 0001
ACM Trans. Priv. Secur.2
2022 STAR: Secret Sharing for Private Threshold Aggregation Reporting
abstract
Threshold aggregation reporting systems promise a practical, privacy-preserving solution for developers to learn how their applications are used in-the-wild. Unfortunately, proposed systems to date prove impractical for wide scale adoption, suffering from a combination of requiring: i) prohibitive trust assumptions; ii) high computation costs; or iii) massive user bases. As a result, adoption of truly-private approaches has been limited to only a small number of enormous (and enormously costly) projects.
Alex Davidson, Peter Snyder, E. B. Quirk, Joseph Genereux, Benjamin Livshits, Hamed Haddadi 0001
CCS6
2022 Tango or square dance?: how tightly should we integrate network functionality in browsers?
abstract
The question at which layer network functionality is presented or abstracted remains a research challenge. Traditionally, network functionality was either placed into the core network, middleboxes, or into the operating system - but recent developments have expanded the design space to directly introduce functionality into the application (and in particular into the browser) as a way to expose it to the user.
Alex Davidson, Matthias Frei, Martin Gartner, Hamed Haddadi 0001, Adrian Perrig, Jordi Subirà Nieto, Philipp Winter, François Wirz
HotNets4
2022 Re-architecting Traffic Analysis with Neural Network Interface Cards
Giuseppe Siracusano, Salvator Galea, Davide Sanvito, Mohammad Malekzadeh, Gianni Antichi, Paolo Costa, Hamed Haddadi 0001, Roberto Bifulco
NSDI7
2022 BatteryLab: A Collaborative Platform for Power Monitoring - https: //batterylab.dev
Matteo Varvello, Kleomenis Katevas, Mihai Plesa, Hamed Haddadi 0001, Fabián E. Bustamante, Benjamin Livshits
PAM4
2022 Blocked or Broken? Automatically Detecting When Privacy Interventions Break Websites
abstract
A core problem in the development and maintenance of crowdsourced filter lists is that their maintainers cannot confidently predict whether (and where) a new filter list rule will break websites. The enormity of the Web prevents filter list authors from broadly understanding the compatibility impact of a new blocking rule before shipping it to millions of users. This severely limits the benefits of filter-list-based content blocking: filter lists are both overly conservative (i.e. rules are tailored narrowly to reduce the risk of breaking things) and error-prone (i.e. blocking tools still break large numbers of sites). To scale to the size and scope of the Web, filter list authors need something better than the current status quo of user reports and manual review, to stop breakage before it has a chance to make it to end users. In this work, we design and implement the first automated system for predicting when a filter list rule breaks a website. We build a classifier, trained on a dataset generated by a combination of compatibility data extracted from the EasyList filter project and novel browser instrumentation, and find that our classifier is accurate to practical levels (AUC 0.88). Our open-source system requires no human interaction when assessing the compatibility risk of a proposed privacy intervention. We also present the 40 page behaviors that most predict breakage in observed websites.
Peter Snyder, Moritz Haller, Benjamin Livshits, Deian Stefan, Hamed Haddadi 0001
Proc. Priv. Enhancing Technol.6
2021 Configurable Privacy-Preserving Automatic Speech Recognition
abstract
Voice assistive technologies have given rise to far-reaching privacy and security concerns. In this paper we investigate whether modular automatic speech recognition (ASR) can improve privacy in voice assistive systems by combining independently trained separation, recognition, and discretization modules to design configurable privacy-preserving ASR systems. We evaluate privacy concerns and the effects of applying various state-of-the-art techniques at each stage of the system, and report results using task-specific metrics (i.e. WER, ABX, and accuracy). We show that overlapping speech inputs to ASR systems present further privacy concerns, and how these may be mitigated using speech separation and optimization techniques. Our discretization module is shown to minimize paralinguistics privacy leakage from ASR acoustic models to levels commensurate with random guessing. We show that voice privacy can be configurable, and argue this presents new opportunities for privacy-preserving applications incorporating ASR.
Ranya Aloufi, Hamed Haddadi 0001, David Boyle 0001
Interspeech2
2021 PPFL: privacy-preserving federated learning with trusted execution environments
abstract
We propose and implement a Privacy-preserving Federated Learning ( PPFL ) framework for mobile systems to limit privacy leakages in federated learning. Leveraging the widespread presence of Trusted Execution Environments (TEEs) in high-end and mobile devices, we utilize TEEs on clients for local training, and on servers for secure aggregation, so that model/gradient updates are hidden from adversaries. Challenged by the limited memory size of current TEEs, we leverage greedy layer-wise training to train each model's layer inside the trusted area until its convergence. The performance evaluation of our implementation shows that PPFL can significantly improve privacy while incurring small system overheads at the client-side. In particular, PPFL can successfully defend the trained model against data reconstruction, property inference, and membership inference attacks. Furthermore, it can achieve comparable model utility with fewer communication rounds (0.54×) and a similar amount of network traffic (1.002×) compared to the standard federated learning of a complete model. This is achieved while only introducing up to ~15% CPU time, ~18% memory usage, and ~21% energy consumption overhead in PPFL's client-side.
Fan Mo 0004, Hamed Haddadi 0001, Kleomenis Katevas, Eduard Marin, Diego Perino, Nicolas Kourtellis
MobiSys2
2021 Stronger Privacy for Federated Collaborative Filtering With Implicit Feedback
abstract
Recommender systems are commonly trained on centrally-collected user interaction data like views or clicks. This practice however raises serious privacy concerns regarding the recommender’s collection and handling of potentially sensitive data. Several privacy-aware recommender systems have been proposed in recent literature, but comparatively little attention has been given to systems at the intersection of implicit feedback and privacy. To address this shortcoming, we propose a practical federated recommender system for implicit data under user-level local differential privacy (LDP). The privacy-utility trade-off is controlled by parameters ϵ and k, regulating the per-update privacy budget and the number of ϵ-LDP gradient updates sent by each user, respectively. To further protect the user’s privacy, we introduce a proxy network to reduce the fingerprinting surface by anonymizing and shuffling the reports before forwarding them to the recommender. We empirically demonstrate the effectiveness of our framework on the MovieLens dataset, achieving up to Hit Ratio with K=10 ([email protected]) 0.68 on 50,000 users with 5,000 items. Even on the full dataset, we show that it is possible to achieve reasonable utility with [email protected]>0.5 without compromising user privacy.
Lorenzo Minto, Moritz Haller, Benjamin Livshits, Hamed Haddadi 0001
RecSys4
2021 Blocking Without Breaking: Identification and Mitigation of Non-Essential IoT Traffic
abstract
Abstract Despite the prevalence of Internet of Things (IoT) devices, there is little information about the purpose and risks of the Internet traffic these devices generate, and consumers have limited options for controlling those risks. A key open question is whether one can mitigate these risks by automatically blocking some of the Internet connections from IoT devices, without rendering the devices inoperable. In this paper, we address this question by developing a rigorous methodology that relies on automated IoT-device experimentation to reveal which network connections (and the information they expose) are essential, and which are not. We further develop strategies to automatically classify network traffic destinations as either required (i.e., their traffic is essential for devices to work properly) or not, hence allowing firewall rules to block traffic sent to non-required destinations without breaking the functionality of the device. We find that indeed 16 among the 31 devices we tested have at least one blockable non-required destination, with the maximum number of blockable destinations for a device being 11. We further analyze the destination of network traffic and find that all third parties observed in our experiments are blockable, while first and support parties are neither uniformly required or non-required. Finally, we demonstrate the limitations of existing blocklists on IoT traffic, propose a set of guidelines for automatically limiting non-essential IoT traffic, and we develop a prototype system that implements these guidelines.
Anna Maria Mandalari, Daniel J. Dubois, Roman Kolcun, Muhammad Talha Paracha, Hamed Haddadi 0001, David R. Choffnes
Proc. Priv. Enhancing Technol.5
2020 A Haystack Full of Needles: Scalable Detection of IoT Devices in the Wild
abstract
Consumer Internet of Things (IoT) devices are extremely popular, providing users with rich and diverse functionalities, from voice assistants to home appliances. These functionalities often come with significant privacy and security risks, with notable recent large-scale coordinated global attacks disrupting large service providers. Thus, an important first step to address these risks is to know what IoT devices are where in a network. While some limited solutions exist, a key question is whether device discovery can be done by Internet service providers that only see sampled flow statistics. In particular, it is challenging for an ISP to efficiently and effectively track and trace activity from IoT devices deployed by its millions of subscribers---all with sampled network data.
Said Jawad Saidi, Anna Maria Mandalari, Roman Kolcun, Hamed Haddadi 0001, Daniel J. Dubois, David R. Choffnes, Georgios Smaragdakis, Anja Feldmann
Internet Measurement Conference4
2020 DarkneTZ: towards model privacy at the edge using trusted execution environments
abstract
We present DarkneTZ, a framework that uses an edge device's Trusted Execution Environment (TEE) in conjunction with model partitioning to limit the attack surface against Deep Neural Networks (DNNs). Increasingly, edge devices (smartphones and consumer IoT devices) are equipped with pre-trained DNNs for a variety of applications. This trend comes with privacy risks as models can leak information about their training data through effective membership inference attacks (MIAs).
Fan Mo 0004, Ali Shahin Shamsabadi, Kleomenis Katevas, Soteris Demetriou, Ilias Leontiadis, Andrea Cavallaro, Hamed Haddadi 0001
MobiSys7
2020 Towards identifying IoT traffic anomalies on the home gateway: poster abstract
abstract
The number of IoT devices continues to grow despite the alarming rate of identification of security and privacy issues. There is widespread concern that development of IoT devices is performed without sufficient attention paid to security and privacy issues. Consequently, networks have a higher probability of incorporating vulnerable IoT devices that may be easy to compromise to launch cyber attacks. Inclusion of IoT devices paves the way for a new category of anomalies to be introduced to networks. Traditional anomaly detection techniques (e.g., semi-supervised and signature-based methods), however, are likely inefficient in detecting IoT-based anomalies. This is because these techniques require static signatures of known attacks, specialized hardware, or full packet inspection. They are also expensive, and may be inaccurate or unscalable. Vulnerable IoT devices can be used to perform destructive attacks or invade privacy. The ability to find anomalies in IoT traffic has the potential to assist with early detection and deployment of countermeasures to thwart such attacks. Thus, new techniques for detecting infected IoT devices are needed to mitigate the associated security and privacy risks. In this research, we investigate the possibility to identify IoT traffic using a combination of behavioural profile, predefined blocklist and device fingerprint. Such a system may be able to detect anomalous and/or malicious devices and/or traffic reliably and quickly. Initial results show that for our implementation of such a system, IoT traffic can be identified using device behaviour profile, fingerprint, and contacted destinations. This work takes the first step towards designing and evaluating iDetector, a framework that can detect anomalous behaviour within IoT networks. In our experiments, iDetector was able to correctly identify 80--90% of all captured traffic traversing a home gateway.
Eman Maali, David Boyle 0001, Hamed Haddadi 0001
SenSys3
2020 A Hybrid Deep Learning Architecture for Privacy-Preserving Mobile Analytics
abstract
Internet-of-Things (IoT) devices and applications are being deployed in our homes and workplaces. These devices often rely on continuous data collection to feed machine learning models. However, this approach introduces several privacy and efficiency challenges, as the service operator can perform unwanted inferences on the available data. Recently, advances in edge processing have paved the way for more efficient, and private, data processing at the source for simple tasks and lighter models, though they remain a challenge for larger and more complicated models. In this article, we present a hybrid approach for breaking down large, complex deep neural networks for cooperative, and privacy-preserving analytics. To this end, instead of performing the whole operation on the cloud, we let an IoT device to run the initial layers of the neural network, and then send the output to the cloud to feed the remaining layers and produce the final result. In order to ensure that the user's device contains no extra information except what is necessary for the main task and preventing any secondary inference on the data, we introduce Siamese fine-tuning. We evaluate the privacy benefits of this approach based on the information exposed to the cloud service. We also assess the local inference cost of different layers on a modern handset. Our evaluations show that by using Siamese fine-tuning and at a small processing cost, we can greatly reduce the level of unnecessary, potentially sensitive information in the personal data, thus achieving the desired tradeoff between utility, privacy, and performance.
Seyed Ali Ossia, Ali Shahin Shamsabadi, Sina Sajadmanesh, Ali Taheri, Kleomenis Katevas, Hamid R. Rabiee 0001, Nicholas D. Lane, Hamed Haddadi 0001
IEEE Internet Things J.8
2020 Privacy and utility preserving sensor-data transformations
Mohammad Malekzadeh, Richard G. Clegg, Andrea Cavallaro, Hamed Haddadi 0001
Pervasive Mob. Comput.4
2020 When Speakers Are All Ears: Characterizing Misactivations of IoT Smart Speakers
abstract
Abstract Internet-connected voice-controlled speakers, also known as smart speakers, are increasingly popular due to their convenience for everyday tasks such as asking about the weather forecast or playing music. However, such convenience comes with privacy risks: smart speakers need to constantly listen in order to activate when the “wake word” is spoken, and are known to transmit audio from their environment and record it on cloud servers. In particular, this paper focuses on the privacy risk from smart speaker misactivations, i.e., when they activate, transmit, and/or record audio from their environment when the wake word is not spoken. To enable repeatable, scalable experiments for exposing smart speakers to conversations that do not contain wake words, we turn to playing audio from popular TV shows from diverse genres. After playing two rounds of 134 hours of content from 12 TV shows near popular smart speakers in both the US and in the UK, we observed cases of 0.95 misactivations per hour, or 1.43 times for every 10,000 words spoken, with some devices having 10% of their misactivation durations lasting at least 10 seconds. We characterize the sources of such misactivations and their implications for consumers, and discuss potential mitigations.
Daniel J. Dubois, Roman Kolcun, Anna Maria Mandalari, Muhammad Talha Paracha, David R. Choffnes, Hamed Haddadi 0001
Proc. Priv. Enhancing Technol.6
2020 PrivEdge: From Local to Distributed Private Training and Prediction
abstract
Machine Learning as a Service (MLaaS) operators provide model training and prediction on the cloud. MLaaS applications often rely on centralised collection and aggregation of user data, which could lead to significant privacy concerns when dealing with sensitive personal data. To address this problem, we propose PrivEdge, a technique for privacy-preserving MLaaS that safeguards the privacy of users who provide their data for training, as well as users who use the prediction service. With PrivEdge, each user independently uses their private data to locally train a one-class reconstructive adversarial network that succinctly represents their training data. As sending the model parameters to the service provider in the clear would reveal private information, PrivEdge secret-shares the parameters among two non-colluding MLaaS providers, to then provide cryptographically private prediction services through secure multi-party computation techniques. We quantify the benefits of PrivEdge and compare its performance with state-of-the-art centralised architectures on three privacy-sensitive image-based tasks: individual identification, writer identification, and handwritten letter recognition. Experimental results show that PrivEdge has high precision and recall in preserving privacy, as well as in distinguishing between private and non-private images. Moreover, we show the robustness of PrivEdge to image compression and biased training data. The source code is available at https://github.com/smartcameras/PrivEdge.
Ali Shahin Shamsabadi, Adrià Gascón, Hamed Haddadi 0001, Andrea Cavallaro
IEEE Trans. Inf. Forensics Secur.3
2020 Deep Private-Feature Extraction
abstract
We present and evaluate Deep Private-Feature Extractor (DPFE), a deep model which is trained and evaluated based on information theoretic constraints. Using the selective exchange of information between a user's device and a service provider, DPFE enables the user to prevent certain sensitive information from being shared with a service provider, while allowing them to extract approved information using their model. We introduce and utilize the log-rank privacy, a novel measure to assess the effectiveness of DPFE in removing sensitive information and compare different models based on their accuracy-privacy trade-off. We then implement and evaluate the performance of DPFEon smartphones to understand its complexity, resource demands, and efficiency trade-offs. Our results on benchmark image datasets demonstrate that under moderate resource utilization, DPFE can achieve high accuracy for primary tasks while preserving the privacy of sensitive information.
Seyed Ali Ossia, Ali Taheri, Ali Shahin Shamsabadi, Kleomenis Katevas, Hamed Haddadi 0001, Hamid R. Rabiee 0001
IEEE Trans. Knowl. Data Eng.5
2019 Poster: Towards Characterizing and Limiting Information Exposure in DNN Layers
abstract
Pre-trained Deep Neural Network (DNN) models are increasingly used in smartphones and other user devices to enable prediction services, leading to potential disclosures of (sensitive) information from training data captured inside these models. Based on the concept of generalization error, we propose a framework to measure the amount of sensitive information memorized in each layer of a DNN. Our results show that, when considered individually, the last layers encode a larger amount of information from the training data compared to the first layers. We find that the same DNN architecture trained with different datasets has similar exposure per layer. We evaluate an architecture to protect the most sensitive layers within an on-device Trusted Execution Environment (TEE) against potential white-box membership inference attacks without the significant computational overhead.
Fan Mo 0004, Ali Shahin Shamsabadi, Kleomenis Katevas, Andrea Cavallaro, Hamed Haddadi 0001
CCS5
2019 BatteryLab, A Distributed Power Monitoring Platform For Mobile Devices
abstract
Recent advances in cloud computing have simplified the way that both software development and testing are performed. Unfortunately, this is not true for battery testing for which state of the art test-beds simply consist of one phone attached to a power meter. These test-beds have limited resources, access, and are overall hard to maintain; for these reasons, they often sit idle with no experiment to run. In this paper, we propose to share existing battery testing setups and build BatteryLab, a distributed platform for battery measurements. Our vision is to transform independent battery testing setups into vantage points of a planetary-scale measurement platform offering heterogeneous devices and testing conditions. In the paper, we design and deploy a combination of hardware and software solutions to enable BatteryLab's vision. We then preliminarily evaluate BatteryLab's accuracy of battery reporting, along with some system benchmarking. We also demonstrate how BatteryLab can be used by researchers to investigate a simple research question.
Matteo Varvello, Kleomenis Katevas, Mihai Plesa, Hamed Haddadi 0001, Benjamin Livshits
HotNets4
2019 Information Exposure From Consumer IoT Devices: A Multidimensional, Network-Informed Measurement Approach
abstract
Internet of Things (IoT) devices are increasingly found in everyday homes, providing useful functionality for devices such as TVs, smart speakers, and video doorbells. Along with their benefits come potential privacy risks, since these devices can communicate information about their users to other parties over the Internet. However, understanding these risks in depth and at scale is difficult due to heterogeneity in devices' user interfaces, protocols, and functionality.
Daniel J. Dubois, David R. Choffnes, Anna Maria Mandalari, Roman Kolcun, Hamed Haddadi 0001
Internet Measurement Conference6
2019 Privacy Against Brute-Force Inference Attacks
abstract
Privacy-preserving data release is about disclosing information about useful data while retaining the privacy of sensitive data. Assuming that the sensitive data is threatened by a brute-force adversary, we define Guessing Leakage as a measure of privacy, based on the concept of guessing. After investigating the properties of this measure, we derive the optimal utility-privacy trade-off via a linear program with any f-information adopted as the utility measure, and show that the optimal utility is a concave and piece-wise linear function of the privacy-leakage budget.
Seyed Ali Ossia, Borzoo Rassouli, Hamed Haddadi 0001, Hamid R. Rabiee 0001, Deniz Gündüz
ISIT3
2019 Privacy preserving speech analysis using emotion filtering at the edge: poster abstract
abstract
Voice controlled devices and services are commonplace in consumer IoT. Cloud-based analysis services extract information from voice input using speech recognition techniques. Services providers can build detailed profiles of users' demographics, preferences and emotional states, etc., and may therefore significantly compromise privacy. To address this problem, a privacy-preserving intermediate layer between users and cloud services is proposed to sanitize voice input directly at edge devices by generating neutralized signals for forwarding. We show that a trained model, based on CycleGAN and deployed on a Raspberry Pi, enables identification and removal of sensitive emotional state information by ~91%, with minimal losses to speech recognition accuracy.
Ranya Aloufi, Hamed Haddadi 0001, David Boyle 0001
SenSys2
2019 "Sensing" the IoT network: Ethical capture of domestic IoT network traffic: poster abstract
abstract
As more and more devices are connected to the Internet-of-Things, often made by non-specialist companies or short-lived startups, the likelihood that these devices will be hacked and used for nefarious activity online increases. We seek to support non-expert users in managing the network behaviour of their IoT devices, and assisting them in handling the cases where those devices are hacked. To do so, we wish to enable anomaly detection at the network level, determining when a device starts behaving unusually. This requires capturing data about how devices behave in a diverse range of real deployments, not just lab environments.
Diana Andreea Popescu, Vadim Safronov, Poonam Yadav, Roman Kolcun, Anna Maria Mandalari, Hamed Haddadi 0001, Derek McAuley, Richard Mortier
SenSys6
2019 BatteryLab, a distributed power monitoring platform for mobile devices: demo abstract
abstract
There has been a growing interest in measuring and optimizing the power efficiency of mobile apps. Traditional power evaluations rely either on inaccurate software-based solutions or on ad-hoc testbeds composed of a power meter and a mobile device. This demonstration presents BatteryLab, our solution to share existing battery testing setups to build a distributed platform for battery measurements. Our vision is to transform independent battery testing setups into vantage points of a planetary-scale measurement platform offering heterogeneous devices and testing conditions. We demonstrate BatteryLab functionalities by investigating the energy efficiency of popular websites when loaded via both Android and iOS browsers. Our demonstration is also live at https://batterylab.dev/.
Matteo Varvello, Kleomenis Katevas, Wei Hang 0003, Mihai Plesa, Hamed Haddadi 0001, Fabián E. Bustamante, Benjamin Livshits
SenSys5
2018 What to Put on the User: Sensing Technologies for Studies and Physiology Aware Systems
abstract
Fitness trackers not just provide easy means to acquire physiological data in real-world environments due to affordable sensing technologies, they further offer opportunities for physiology-aware applications and studies in HCI; however, their performance is not well understood. In this paper, we report findings on the quality of 3 sensing technologies: PPG-based wrist trackers (Apple Watch, Microsoft Band 2), an ECG-belt (Polar H7) and reference device with stick-on ECG electrodes (Nexus 10). We collected physiological (heart rate, electrodermal activity, skin temperature) and subjective data from 21 participants performing combinations of physical activity and stressful tasks. Our empirical research indicates that wrist devices provide a good sensing performance in stationary settings. However, they lack accuracy when participants are mobile or if tasks require physical activity. Based on our findings, we suggest a textitDesign Space for Wearables in Research Settings and reflected on the appropriateness of the investigated technologies in research contexts.
Katrin Hänsel, Romina Poguntke, Hamed Haddadi 0001, Akram Alomainy, Albrecht Schmidt 0001
CHI3
2018 Distributed One-Class Learning
abstract
We propose a cloud-based filter trained to block third parties from uploading privacy-sensitive images of others to online social media. The proposed filter uses Distributed One-Class Learning, which decomposes the cloud-based filter into multiple one-class classifiers. Each one-class classifier captures the properties of a class of privacy-sensitive images with an autoencoder. The multi-class filter is then reconstructed by combining the parameters of the one-class autoen-coders. The training takes place on edge devices (e.g. smartphones) and therefore users do not need to upload their private and/or sensitive images to the cloud. A major advantage of the proposed filter over existing distributed learning approaches is that users cannot access, even indirectly, the parameters of other users. Moreover, the filter can cope with the imbalanced and complex distribution of the image content and the independent probability of addition of new users. We evaluate the performance of the proposed distributed filter using the exemplar task of blocking a user from sharing privacy-sensitive images of other users. In particular, we validate the behavior of the proposed multi-class filter with non-privacy-sensitive images, the accuracy when the number of classes increases, and the robustness to attacks when an adversary user has access to privacy-sensitive images of other users.
Ali Shahin Shamsabadi, Hamed Haddadi 0001, Andrea Cavallaro
ICIP2
2018 Special theme on privacy and the Internet of things
Alan Chamberlain, Andy Crabtree, Hamed Haddadi 0001, Richard Mortier
Pers. Ubiquitous Comput.3
2017 Demo: AWSense: A Framework for Collecting Sensing Data from the Apple Watch
abstract
We present a framework - AWSense - for eased sensing data collection from an Apple Watch wearable device. The framework eases the access, transmission and export of sensing data from the device. This data comprises: heart rate, raw acceleration, and computed device motion. In our demo, we present sample applications built on top of this framework to show its capabilities, of real-time presentation and recording of the sensing data.
Katrin Hänsel, Hamed Haddadi 0001, Akram Alomainy
MobiSys2
2017 Demo: Detecting Group Formations using iBeacon Technology
abstract
Researchers from different disciplines have examined crowd behavior in the past by employing a variety of methods including ethnographic studies, computer vision techniques and manual annotation based data analysis. However, because of the inherent difficulties in collecting, processing and analyzing the data, it is difficult to obtain large data sets for study. In this work we present a system for detecting stationary interactions inside crowds, depending entirely on the sensors available in a modern smartphone device such as Bluetooth Smart (BLE) and Accelerometer. By utilizing Apple's iBeaconTM implementation of Bluetooth Smart using SensingKit1, our open-source multi-platform mobile sensing framework [1], we are able to detect the proximity of users carrying a smartphone in their pocket. We then use an algorithm based on graph theory to predict group interactions inside the crowd. Previous work in this area has been limited to the detection of interactions between only two people and therefore our approach goes beyond current state of the art in its ability to detect group formations with more than two people involved. Our approach is particularly beneficial to the design and implementation of crowd behavior analytics, design of influence strategies, and algorithms for crowd reconfiguration.
Kleomenis Katevas, Laurissa N. Tokarchuk, Hamed Haddadi 0001, Richard G. Clegg
MobiSys3
2016 A first look at user activity on tinder
abstract
Mobile dating apps have become a popular means to meet potential partners. Although several exist, one recent addition stands out amongst all others. Tinder presents its users with pictures of people geographically nearby, whom they can either like or dislike based on first impressions. If two users like each other, they are allowed to initiate a conversation via the chat feature. In this paper we use a set of curated profiles to explore the behaviour of men and women in Tinder. We reveal differences between the way men and women interact with the app, highlighting the strategies employed. Women attain large numbers of matches rapidly, whilst men only slowly accumulate matches. Most notably, our results indicate that a little effort in grooming profiles, especially for male users, goes a long way in attracting attention.
Gareth Tyson, Vasile Claudiu Perta, Hamed Haddadi 0001, Michael C. Seto
ASONAM3
2016 Fetishizing Food in Digital Age: #foodporn Around the World
Yelena Mejova, Sofiane Abbar, Hamed Haddadi 0001
ICWSM3
2016 SensingKit: Evaluating the Sensor Power Consumption in iOS Devices
abstract
Today's smartphones come equipped with a range of advanced sensors capable of sensing motion, orientation, audio as well as environmental data with high accuracy. With the existence of application distribution channels such as the Apple App Store and the Google Play Store, researchers can distribute applications and collect large scale data in ways that previously were not possible. Motivated by the lack of a universal, multi-platform sensing library, in this work we present the design and implementation of SensingKit, an open-source continuous sensing system that supports both iOS and Android mobile devices. One of the unique features of SensingKit is the support of the latest beacon technologies based on Bluetooth Smart (BLE), such as iBeacon and Eddystone. We evaluate and compare the power consumption of each supported sensor individually, using an iPhone 5S device running on iOS 9. We believe that this platform will be beneficial to all researchers and developers who plan to use mobile sensing technology in large-scale experiments.
Kleomenis Katevas, Hamed Haddadi 0001, Laurissa N. Tokarchuk
Intelligent Environments2
2016 Tracking Personal Identifiers Across the Web
Marjan Falahrastegar, Hamed Haddadi 0001, Steve Uhlig, Richard Mortier
PAM2
2016 Privacy-Aware Infrastructure for Managing Personal Data
abstract
In recent times, we have seen a proliferation of personal data. This can be attributed not just to a larger proportion of our lives moving online, but also through the rise of ubiquitous sensing through mobile and IoT devices. Alongside this surge, concerns over privacy, trust, and security are expressed more and more as different parties attempt to take advantage of this rich assortment of data.
Yousef Amar, Hamed Haddadi 0001, Richard Mortier
SIGCOMM2
2016 Enabling the new economic actor: data protection, the digital economy, and the Databox
abstract
This paper offers a sociological perspective on data protection regulation and its relevance to design. From this perspective, proposed regulation in Europe and the USA seeks to create a new economic actor—the consumer as personal data trader—through new legal frameworks that shift the locus of agency and control in data processing towards the individual consumer or “data subject”. The sociological perspective on proposed data regulation recognises the reflexive relationship between law and the social order, and the commensurate needs to balance the demand for compliance with the design of computational tools that enable this new economic actor. We present the Databox model as a means of providing data protection and allowing the individual to exploit personal data to become an active player in the emerging data economy.
Andy Crabtree, Tom Lodge, James A. Colley, Christopher Greenhalgh, Richard Mortier, Hamed Haddadi 0001
Pers. Ubiquitous Comput.6
2015 A Glance through the VPN Looking Glass: IPv6 Leakage and DNS Hijacking in Commercial VPN clients
abstract
Abstract Commercial Virtual Private Network (VPN) services have become a popular and convenient technology for users seeking privacy and anonymity. They have been applied to a wide range of use cases, with commercial providers often making bold claims regarding their ability to fulfil each of these needs, e.g., censorship circumvention, anonymity and protection from monitoring and tracking. However, as of yet, the claims made by these providers have not received a sufficiently detailed scrutiny. This paper thus investigates the claims of privacy and anonymity in commercial VPN services. We analyse 14 of the most popular ones, inspecting their internals and their infrastructures. Despite being a known issue, our experimental study reveals that the majority of VPN services suffer from IPv6 traffic leakage. The work is extended by developing more sophisticated DNS hijacking attacks that allow all traffic to be transparently captured.We conclude discussing a range of best practices and countermeasures that can address these vulnerabilities
Vasile Claudiu Perta, Marco Valerio Barbera, Gareth Tyson, Hamed Haddadi 0001, Alessandro Mei
Proc. Priv. Enhancing Technol.4
2014 Poster: SensingKit: a multi-platform mobile sensing framework for large-scale experiments
abstract
With the rapid rise in variety of available smartphones today and their rich sensing capabilities, there is an increasing interest in using mobile sensing in large-scale experiments and commercial applications. Motivated by the lack of a universal, multi-platform library, in this paper we present SensingKit, an efficient, open-source, client-server system that supports both iOS and Android mobile devices. SensingKit is capable of continuous sensing the device's motion (Accelerometer, Gyroscope, Magnetometer), location (GPS) and proximity to other smartphones (Bluetooth Smart). The data are temporarily saved to the device's memory and transmitted to a server for further analysis over any Internet connection. We believe that this platform will be beneficial to all researchers and developers who need to perform mobile sensing in their applications and experiments.
Kleomenis Katevas, Hamed Haddadi 0001, Laurissa N. Tokarchuk
MobiCom2
2012 Breaking for commercials: characterizing mobile advertising
abstract
Mobile phones and tablets can be considered as the first incarnation of the post-PC era. Their explosive adoption rate has been driven by a number of factors, with the most signifcant influence being applications (apps) and app markets. Individuals and organizations are able to develop and publish apps, and the most popular form of monetization is mobile advertising.
Narseo Vallina-Rodriguez, Jay Shah, Alessandro Finamore, Yan Grunenberger, Konstantina Papagiannaki, Hamed Haddadi 0001, Jon Crowcroft
Internet Measurement Conference6
2012 The World of Connections and Information Flow in Twitter
abstract
Information propagation in online social networks like Twitter is unique in that word-of-mouth propagation and traditional media sources coexist. We collect a large amount of data from Twitter to compare the relative roles different types of users play in information flow. Using empirical data on the spread of news about major international headlines as well as minor topics, we investigate the relative roles of three types of information spreaders: 1) mass media sources like BBC; 2) grassroots, consisting of ordinary users; and 3) evangelists, consisting of opinion leaders, politicians, celebrities, and local businesses. Mass media sources play a vital role in reaching the majority of the audience in any major topics. Evangelists, however, introduce both major and minor topics to audiences who are further away from the core of the network and would otherwise be unreachable. Grassroots users are relatively passive in helping spread the news, although they account for the 98% of the network. Our results bring insights into what contributes to rapid information propagation at different levels of topic popularity, which we believe are useful to the designers of social search and recommendation engines.
Meeyoung Cha, Fabrício Benevenuto, Hamed Haddadi 0001, Krishna P. Gummadi
IEEE Trans. Syst. Man Cybern. Part A3
2011 Discriminating graphs through spectral projections
Damien Fay, Hamed Haddadi 0001, Steve Uhlig, Liam Kilmartin, Andrew W. Moore 0002, Jérôme Kunegis, Marios Iliofotou
Comput. Networks2
2010 Measuring User Influence in Twitter: The Million Follower Fallacy
Meeyoung Cha, Hamed Haddadi 0001, Fabrício Benevenuto, Krishna P. Gummadi
ICWSM2
2010 Weighted spectral distribution for internet topology analysis: theory and applications
Damien Fay, Hamed Haddadi 0001, Andrew Thomason 0001, Andrew W. Moore 0002, Richard Mortier, Almerima Jamakovic, Steve Uhlig, Miguel Rio
IEEE/ACM Trans. Netw.2
2009 Serving Ads from localhost for Performance, Privacy, and Profit
Saikat Guha 0002, Alexey Reznichenko, Kevin Tang, Hamed Haddadi 0001, Paul Francis
HotNets4
2009 Challenges in the capture and dissemination of measurements from high-speed networks
abstract
The production of a large-scale monitoring system for a high-speed network leads to a number of challenges. These challenges are not purely technical but also socio-political and legal. The number of stakeholders in such monitoring activity is large including the network operators, the users, the equipment manufacturers and, of course, the monitoring researchers. The MASTS project (measurement at all scales in time and space) was created to instrument the high-speed JANET Lightpath network and has been extended to incorporate other paths supported by JANET(UK). Challenges the project has faced included: simple access to the network; legal issues involved in the storage and dissemination of the captured information, which may be personal; the volume of data captured and the rate at which these data appear at store. To this end, the MASTS system will have established four monitoring points each capturing packets on a high-speed link. Traffic header data will be continuously collected, anonymised, indexed, stored and made available to the research community. A legal framework for the capture and storage of network measurement data has been developed which allows the anonymised IP traces to be used for research purposes.
Richard G. Clegg, Mark S. Withall, Andrew W. Moore 0002, Iain Phillips 0002, David J. Parish, Miguel Rio, Raul Landa, Hamed Haddadi 0001, Konstantinos G. Kyriakopoulos, J. Auge, R. Clayton, D. Salmon
IET Commun.8