EDBT 2026 Demo / reviewers in the wild / expert
Nicholas D. Lane
dblp:03/2663 · also Nicholas Donald Lane
· DBLP profile ↗
132ranked-venue papers
12as first author
56since 2021 · last 2026
0000-0002-2728-8273ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 48 · 33 since 2021Computer networks · 44 · 5 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 24 · 12 since 2021Human-computer interaction and ubiquitous computing · 16 · 7 first-author · 1 since 2021Systems, architecture and hardware · 15 · 11 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Security and privacy · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Editing as Unlearning: Are Knowledge Editing Methods Strong Baselines for Large Language Model Unlearning?abstractLarge language Model (LLM) unlearning, i.e., selectively removing information from LLMs, is vital for responsible model deployment. Differently, LLM knowledge editing aims to modify LLM knowledge instead of removing it. Though editing and unlearning seem to be two distinct tasks, we find there is a tight connection between them. In this paper, we conceptualize unlearning as a special case of editing where information is modified to a refusal or "empty set" response, signifying its removal. This paper thus investigates if knowledge editing techniques are strong baselines for LLM unlearning. We evaluate state-of-the-art (SOTA) editing methods (e.g., ROME, MEMIT, GRACE, WISE, and AlphaEdit) against existing unlearning approaches on pretrained and finetuned knowledge. Results show certain editing methods, notably WISE and AlphaEdit, are effective unlearning baselines, especially for pretrained knowledge, and excel in generating human-aligned refusal answers. To better adapt editing methods for unlearning applications, we propose practical recipes including self-improvement and query merging. The former leverages the LLM's own in-context learning ability to craft a more human-aligned unlearning target, and the latter enables ROME and MEMIT to perform well in unlearning longer sample sequences. We advocate for the unlearning community to adopt SOTA editing methods as baselines and explore unlearning from an editing perspective for more holistic LLM memory control. Zexi Li 0001, Xiangzhu Wang, William F. Shen, Meghdad Kurmanji, Xinchi Qiu, Dongqi Cai 0001, Chao Wu 0001, Nicholas D. Lane |
AAAI | 8 |
| 2026 | Efficient 3D Gaussian Splatting with Axis-Shared Rasterization and Order-independent Transmittanceabstract3D Gaussian Splatting (3DGS) has emerged as a powerful technique for novel view synthesis, combining high-quality reconstruction with efficient rendering. It has been widely adopted in domains such as AR/VR, robotics, and autonomous driving. However, achieving real-time performance on resource-constrained platforms remains challenging due to strict power and area budgets. Prior accelerators improve hardware performance but still overlook key inefficiencies, including insufficient rasterization efficiency, poor sorting scalability, and pipeline imbalance. This paper presents an architecture-algorithm co-design to address these challenges. First, we propose axis-shared rasterization, which precomputes and reuses common terms along the X- and Y-axes, reducing multiply-and-accumulate (MAC) operations by up to 38% while preserving high parallelism. Second, we develop a novel order-independent transmittance method that removes the need for explicit sorting by leveraging a lightweight multilayer perceptron (MLP) to directly approximate the transmittance of each Gaussian, enabling efficient alpha blending with negligible quality loss. Third, we design a unified reconfigurable PE array that supports both rasterization and MLP inference, sustaining high utilization without costly sorting hardware. Our experiments demonstrate that our design preserves rendering quality while achieving a 1.33 to 1.88x speedup over state-of-the-art 3DGS accelerators. Our code is open source at https://github.com/WangZhican/ISCA26_3DGS_Acc. Zhican Wang, Guanghui He 0002, Lingjun Gao, Dantong Liu, Shell Xu Hu, Chen Zhang 0001, Zhuoran Song, Nicholas D. Lane, Hongxiang Fan |
ISCA | 8 |
| 2026 | The Triple Lock: Geometric Obstructions to Partial Information Decomposition
Vivek Kothari, Nicholas D. Lane |
ISIT | 2 |
| 2026 | Keynote - Federated AGI: A Generational Opportunity for Pervasive ComputingabstractCurrent scaling laws indicate that future advances in AI will hinge on access to massive amounts of compute and data. How will we obtain the computing power and data resources required to sustain continued AI progress? I believe all roads lead to federated learning, and approaches of this kind. And, in turn, this change represents the biggest opportunity for pervasive computing in a generation: with the right research priorities and execution, pervasive systems can provide the critical infrastructure for tomorrow’s frontier AI, supplying both compute and data; serving as the missing piece for the AI miracle to continue.The AI world may not realize it yet, but it is ready for this radical change. Today, frontier models are trained on essentially the same web-scraped data and rely on data centers that can only scale through unsustainable capital investment. Instead, pervasive systems of sensors, personal devices, scientific instruments, space satellites, vehicles, and embedded computers offer both a limitless stream of new data and a modular, scalable network of compute. Once resources of this scale and quality are unlocked, today’s AI infrastructure paradigm will not compete. Soon, decentralized and federated techniques built on pervasive systems will be how the strongest LLMs are trained; and in time, how aspirational capabilities like AGI will finally be achieved.In this talk, I will describe why the future of AI will be federated, and describe early solutions from Flower Labs and the Cambridge ML Systems Lab (CaMLSys), which address the technical challenges of this shift and place pervasive computing and communication at the core of how tomorrow’s AI is built. Nicholas D. Lane |
PerCom | 1 |
| 2026 | FwdLLM+: Accelerating Forward-Only FedLLM With Low-Rank PerturbationsabstractFederated Learning (FL) facilitates privacy-preserving fine-tuning of Large Language Models (LLMs) for mobile applications, termed FedLLM. A vital challenge of FedLLM is the tension between LLM complexity and resource constraint of mobile devices. In response to this challenge, we first introduceFwdLLM(our conference version), an innovative FL framework designed to enhance the FedLLM efficiency. The key idea ofFwdLLMis to employ backpropagation (BP)-free training methods, requiring devices only to execute memory-efficient “perturbed inference”. Enabled by mobile NPU acceleration and an expanded array of participating devices,FwdLLMdelivers substantially better wall-clock efficiency than BP-based FedLLM. However,FwdLLMbased on vanilla BP-free optimization theoretically requires more optimization steps to converge. In this work, we further enhanceFwdLLMtoFwdLLM+, which incorporates advanced zeroth-order optimization techniques and low-rank perturbation decomposition to reduce convergence steps. Finally, we conduct extensive experiments on 4 models (ranging from 110 M to 7B) and 8 more datasets, demonstrating thatFwdLLM+achieves up to 151× faster training, a 93$\%$memory reduction compared to vanilla BP-based FedLLM, and superior performance compared toFwdLLM, enabling efficient federated fine-tuning of billion-parameter LLMs on commodity mobile devices. Mengwei Xu 0001, Zhenyan Lu, Wei Liu 0302, Shangguang Wang, Nicholas D. Lane, Qibo Sun, Dongqi Cai 0001 |
IEEE Trans. Mob. Comput. | 6 |
| 2026 | Floe: Federated Specialization for Real-Time LLM-SLM InferenceabstractDeploying large language models (LLMs) in realtime systems is challenging due to their high resource demands and privacy concerns. We propose Floe a hybrid federated learning framework designed for latency-sensitive, resourceconstrained environments. Floeombines a cloud-based blackbox LLM with lightweight small language models (SLMs) on edge devices to enable low-latency, privacy-preserving inference. Personal data and fine-tuning remain on-device, while the cloud LLM contributes general knowledge without exposing proprietary weights. A heterogeneity-aware LoRA adaptation strategy ensures efficient edge deployment across diverse hardware, and a logit-level fusion mechanism enables real-time coordination between edge and cloud models. Experiments demonstrate that Floenhances user privacy and personalization, while significantly improving model performance and reducing inference latency on edge devices under real-time constraints, compared to baseline approaches. Chunlin Tian, Kahou Tam, Yebo Wu, Shuaihang Zhong, Li Li 0064, Nicholas D. Lane, Cheng-Zhong Xu 0001 |
IEEE Trans. Parallel Distributed Syst. | 6 |
| 2025 | Demystifying Small Language Models for Edge DeploymentabstractSmall language models (SLMs) have emerged as a promising solution for deploying resource-constrained devices, such as smartphones and Web of Things. This work presents the first comprehensive study of over 60 SLMs such as Microsoft Phi and Google Gemma that are publicly accessible. Our findings show that state-of-the-art SLMs outperform 7B models in general tasks, proving their practical viability. However, SLMs’ in-context learning capabilities remain limited, and their efficiency has significant optimization potential. We identify key SLM optimization opportunities, including dynamic task-specific routing, model-hardware co-design, and vocabulary/KV cache compression. Overall, we expect the work to reveal an all-sided landscape of SLMs, benefiting the research community across algorithm, model, system, and hardware levels. Zhenyan Lu, Xiang Li 0067, Dongqi Cai 0001, Rongjie Yi, Fangming Liu, Wei Liu 0302, Jian Luan 0001, Nicholas D. Lane, Mengwei Xu 0001 |
ACL (1) | 9 |
| 2025 | Exploring Code Language Models for Automated HLS-based Hardware Generation: Benchmark, Infrastructure and AnalysisabstractRecent advances in code generation have illuminated the potential of employing large language models (LLMs) for general-purpose programming languages such as Python and C++, opening new opportunities for automating software development and enhancing programmer productivity. The potential of LLMs in software programming has sparked significant interest in exploring automated hardware generation and automation. Although preliminary endeavors have been made to adopt LLMs in generating hardware description languages (HDLs) such as Verilog and SystemVerilog, several challenges persist in this direction. First, the volume of available HDL training data is substantially smaller compared to that for software programming languages. Second, the pre-trained LLMs, mainly tailored for software code, tend to produce HDL designs that are more error-prone. Third, the generation of HDL requires a significantly higher number of tokens compared to software programming, leading to inefficiencies in cost and energy consumption. To tackle these challenges, this paper explores leveraging LLMs to generate High-Level Synthesis (HLS)-based hardware design. Although code generation for domain-specific programming languages is not new in the literature, we aim to provide experimental results, insights, benchmarks, and evaluation infrastructure to investigate the suitability of HLS over low-level HDLs for LLM-assisted hardware design generation. To achieve this, we first finetune pretrained models for HLS-based hardware generation, using a collected dataset with text prompts and corresponding reference HLS designs. An LLM-assisted framework is then proposed to automate end-to-end hardware code generation, which also investigates the impact of chain-of-thought and feedback loops promoting techniques on HLS- design generation. Comprehensive experiments demonstrate the effectiveness of our methods. Jiahao Gai, Zhican Wang, Wanru Zhao, Nicholas D. Lane, Hongxiang Fan |
ASP-DAC | 6 |
| 2025 | SparsyFed: Sparse Adaptive Federated LearningabstractSparse training is often adopted in cross-device federated learning (FL) environments where constrained devices collaboratively train a machine learning model on private data by exchanging pseudo-gradients across heterogeneous networks. Although sparse training methods can reduce communication overhead and computational burden in FL, they are often not used in practice for the following key reasons: (1) data heterogeneity makes it harder for clients to reach consensus on sparse models compared to dense ones, requiring longer training; (2) methods for obtaining sparse masks lack adaptivity to accommodate very heterogeneous data distributions, crucial in cross-device FL; and (3) additional hyperparameters are required, which are notably challenging to tune in FL. This paper presents SparsyFed, a practical federated sparse training method that critically addresses the problems above. Previous works have only solved one or two of these challenges at the expense of introducing new trade-offs, such as clients’ consensus on masks versus sparsity pattern adaptivity. We show that SparsyFed simultaneously (1) can produce 95% sparse models, with negligible degradation in accuracy, while only needing a single hyperparameter, (2) achieves a per-round weight regrowth 200 times smaller than previous methods, and (3) allows the sparse masks to adapt to highly heterogeneous data distributions and outperform all baselines under such conditions. Adriano Guastella, Lorenzo Sani, Alex Iacob, Alessio Mora, Paolo Bellavista, Nicholas D. Lane |
ICLR | 6 |
| 2025 | DEPT: Decoupled Embeddings for Pre-training Language ModelsabstractLanguage Model pre-training uses broad data mixtures to enhance performance across domains and languages. However, training on such heterogeneous text corpora requires extensive and expensive efforts. Since these data sources vary significantly in lexical, syntactic, and semantic aspects, they cause negative interference or the ``curse of multilinguality''. To address these challenges we propose a communication-efficient pre-training framework, DEPT. Our method decouples embeddings from the transformer body while simultaneously training the latter on multiple data sources without requiring a shared vocabulary. DEPT can: (1) train robustly and effectively under significant data heterogeneity, (2) minimize token embedding parameters to only what the data source vocabulary requires, while cutting communication costs in direct proportion to both the communication frequency and the reduction in parameters, (3) enhance transformer body plasticity and generalization, improving both average perplexity (up to 20%) and downstream task performance, and (4) enable training with custom optimized vocabularies per data source. We demonstrate DEPT's potential via the first vocabulary-agnostic federated pre-training of billion-scale models, reducing communication costs by orders of magnitude and embedding memory by 4-5x. Alex Iacob, Lorenzo Sani, Meghdad Kurmanji, William F. Shen, Xinchi Qiu, Dongqi Cai 0001, Yan Gao 0016, Nicholas D. Lane |
ICLR | 8 |
| 2025 | Rapid Distributed Fine-tuning of a Segmentation Model Onboard SatellitesabstractSegmentation of Earth observation (EO) satellite data is critical for natural hazard analysis and disaster response. However, processing EO data at ground stations introduces delays due to data transmission bottlenecks and communication windows. Using segmentation models capable of near-real-time data analysis onboard satellites can therefore improve response times. This study presents a proof-of-concept using MobileSAM, a lightweight, pre-trained segmentation model, onboard Unibap iX10-100 satellite hardware. We demonstrate the segmentation of water bodies from Sentinel-2 satellite imagery and integrate MobileSAM with PASEOS, an open-source Python module that simulates satellite operations. This integration allows us to evaluate MobileSAM’s performance under simulated conditions of a satellite constellation. Our research investigates the potential of fine-tuning MobileSAM in a decentralised way onboard multiple satellites in rapid response to a disaster. Our findings show that MobileSAM can be rapidly fine-tuned and benefits from decentralised learning, considering the constraints imposed by the simulated orbital environment. We observe improvements in segmentation performance with minimal training data and fast fine-tuning when satellites frequently communicate model updates. This study contributes to the field of onboard AI by emphasising the benefits of decentralized learning and fine-tuning pre-trained models for rapid response scenarios. Our work builds on recent related research at a critical time; as extreme weather events increase in frequency and magnitude, rapid response with onboard data analysis is essential. Meghan Plumridge, Rasmus Maråk, Chiara Ceccobello, Pablo Gómez, Gabriele Meoni, Filip Svoboda, Nicholas D. Lane |
IPAS | 7 |
| 2025 | FedGuCci: Making Local Models More Connected in Landscape for Federated LearningabstractFederated learning (FL) involves multiple heterogeneous clients collaboratively training a global model via iterative local updates and model fusion.The generalization of FL's global model has a large gap compared with centralized training, which is its bottleneck for broader applications.In this paper, we study and improve FL's generalization through a fundamental "connectivity" perspective, which means how the local models are connected in the parameter region and fused into a generalized global model.The term "connectivity" is derived from linear mode connectivity (LMC), studying the interpolated loss landscape of two different solutions (e.g., modes) of neural networks.Bridging the gap between LMC and FL, in this paper, we leverage fixed anchor models to empirically and theoretically study the transitivity property of connectivity from two models (LMC) to a group of models (model fusion in FL).Based on the findings, we propose FedGuCci(+), improving group connectivity for better generalization.It is shown that our methods can boost the generalization of FL under client heterogeneity across various tasks (4 CV datasets and 6 NLP datasets) and model architectures (e.g., ViTs and PLMs).The code is available here: FedGuCci Codebase. Zexi Li 0001, Zhiqi Li 0004, Didi Zhu, Tao Shen 0002, Tao Lin 0004, Chao Wu 0001, Nicholas D. Lane |
KDD (2) | 8 |
| 2025 | FLUX: Efficient Descriptor-Driven Clustered Federated Learning under Arbitrary Distribution ShiftsabstractFederated Learning (FL) enables collaborative model training across multiple clients while preserving data privacy. Traditional FL methods often use a global model to fit all clients, assuming that clients' data are independent and identically distributed (IID). However, when this assumption does not hold, the global model accuracy may drop significantly, limiting FL applicability in real-world scenarios. To address this gap, we propose FLUX, a novel clustering-based FL (CFL) framework that addresses the four most common types of distribution shifts during both training and test time. To this end, FLUX leverages privacy-preserving client-side descriptor extraction and unsupervised clustering to ensure robust performance and scalability across varying levels and types of distribution shifts. Unlike existing CFL methods addressing non-IID client distribution shifts, FLUX i) does not require any prior knowledge of the types of distribution shifts or the number of client clusters, and ii) supports test-time adaptation, enabling unseen and unlabeled clients to benefit from the most suitable cluster-specific models. Extensive experiments across four standard benchmarks, two real-world datasets and ten state-of-the-art baselines show that FLUX improves performance and stability under diverse distribution shifts—achieving an average accuracy gain of up to 23 percentage points over the best-performing baselines—while maintaining computational and communication overhead comparable to FedAvg. Dario Fenoglio, Mohan Li, Pietro Barbiero, Nicholas D. Lane, Marc Langheinrich, Martin Gjoreski |
NeurIPS | 4 |
| 2025 | FlowerTune: A Cross-Domain Benchmark for Federated Fine-Tuning of Large Language ModelsabstractLarge Language Models (LLMs) have achieved state-of-the-art results across diverse domains, yet their development remains reliant on vast amounts of publicly available data, raising concerns about data scarcity and the lack of access to domain-specific, sensitive information. Federated Learning (FL) presents a compelling framework to address these challenges by enabling decentralized fine-tuning on pre-trained LLMs without sharing raw data. However, the compatibility and performance of pre-trained LLMs in FL settings remain largely under explored. We introduce the FlowerTune LLM Leaderboard, a first-of-its-kind benchmarking suite designed to evaluate federated fine-tuning of LLMs across four diverse domains: general NLP, finance, medical, and coding. Each domain includes federated instruction-tuning datasets and domain-specific evaluation metrics. Our results, obtained through a collaborative, open-source and community-driven approach, provide the first comprehensive comparison across 26 pre-trained LLMs with different aggregation and fine-tuning strategies under federated settings, offering actionable insights into model performance, resource constraints, and domain adaptation. This work lays the foundation for developing privacy-preserving, domain-specialized LLMs for real-world applications. Yan Gao 0016, Massimo Roberto Scamarcia, Javier Fernández-Marqués, Mohammad Naseri, Chong Shen Ng, Dimitris Stripelis, Zexi Li 0001, Tao Shen 0002, Jiamu Bai, Daoyuan Chen, Zikai Zhang 0003, Rui Hu 0005, Inseo Song, Kangyoon Lee, Hong Jia, Ting Dang, Zheyuan Liu 0002, Daniel J. Beutel, Lingjuan Lyu, Nicholas D. Lane |
NeurIPS | 21 |
| 2025 | Position: Bridge the Gaps between Machine Unlearning and AI RegulationabstractThe "right to be forgotten" and the data privacy laws that encode it have motivated machine unlearning since its earliest days. Now, some argue that an inbound wave of artificial intelligence regulations — like the European Union's Artificial Intelligence Act (AIA) — may offer important new use cases for machine unlearning. However, this position paper argues, this opportunity will only be realized if researchers proactively bridge the (sometimes sizable) gaps between machine unlearning's state of the art and its potential applications to AI regulation. To demonstrate this point, we use the AIA as our primary case study. Specifically, we deliver a "state of the union" as regards machine unlearning's current potential (or, in many cases, lack thereof) for aiding compliance with the AIA. This starts with a precise cataloging of the potential applications of machine unlearning to AIA compliance. For each, we flag the technical gaps that exist between the potential application and the state of the art of machine unlearning. Finally, we end with a call to action: for machine learning researchers to solve the open technical questions that could unlock machine unlearning's potential to assist compliance with the AIA — and other AI regulation like it. Bill Marino, Meghdad Kurmanji, Nicholas D. Lane |
NeurIPS | 3 |
| 2025 | LLM Unlearning via Neural Activation RedirectionabstractThe ability to selectively remove knowledge from LLMs is highly desirable. However, existing methods often struggle with balancing unlearning efficacy and retain model utility, and lack controllability at inference time to emulate base model behavior as if it had never seen the unlearned data. In this paper, we propose LUNAR, a novel unlearning method grounded in the Linear Representation Hypothesis and operates by redirecting the representations of unlearned data to activation regions that expresses its inability to answer. We show that contrastive features are not a prerequisite for effective activation redirection, and LUNAR achieves state-of-the-art unlearning performance and superior controllability. Specifically, LUNAR achieves between 2.9x and 11.7x improvement in the combined unlearning efficacy and model utility score (Deviation Score) across various base models and generates coherent, contextually appropriate responses post-unlearning. Moreover, LUNAR effectively reduces parameter updates to a single down-projection matrix, a novel design that significantly enhances efficiency by 20x and robustness. Finally, we demonstrate that LUNAR is robust to white-box adversarial attacks and versatile in real-world scenarios, including handling sequential unlearning requests. William F. Shen, Xinchi Qiu, Meghdad Kurmanji, Alex Iacob, Lorenzo Sani, Nicola Cancedda, Nicholas D. Lane |
NeurIPS | 8 |
| 2025 | PQA: Exploring the Potential of Product Quantization in DNN Hardware AccelerationabstractConventional multiply-accumulate (MAC) operations have long dominated computation time for deep neural networks (DNNs), especially convolutional neural networks (CNNs). Recently, product quantization (PQ) has been applied to these workloads, replacing MACs with memory lookups to pre-computed dot products. To better understand the efficiency tradeoffs of product-quantized DNNs (PQ-DNNs), we create a custom hardware accelerator to parallelize and accelerate nearest-neighbor search and dot-product lookups. Additionally, we perform an empirical study to investigate the efficiency–accuracy tradeoffs of different PQ parameterizations and training methods. We identify PQ configurations that improve performance-per-area for ResNet20 by up to 3.1×, even when compared to a highly optimized conventional DNN accelerator, with similar improvements on two additional compact DNNs. When comparing to recent PQ solutions, we outperform prior work by 4× in terms of performance-per-area with a 0.6% accuracy degradation. Finally, we reduce the bitwidth of PQ operations to investigate the impact on both hardware efficiency and accuracy. With only 2–6-bit precision on three compact DNNs, we were able to maintain DNN accuracy eliminating the need for DSPs. Ahmed F. AbouElhamayed, Angela Cui, Javier Fernández-Marqués, Nicholas D. Lane, Mohamed S. Abdelfattah |
ACM Trans. Reconfigurable Technol. Syst. | 4 |
| 2024 | Breaking Physical and Linguistic Borders: Multilingual Federated Prompt Tuning for Low-Resource LanguagesabstractPretrained large language models (LLMs) have emerged as a cornerstone in modern natural language processing, with their utility expanding to various applications and languages. However, the fine-tuning of multilingual LLMs, particularly for low-resource languages, is fraught with challenges steming from data-sharing restrictions (the physical border) and from the inherent linguistic differences (the linguistic border). These barriers hinder users of various languages, especially those in low-resource regions, from fully benefiting from the advantages of LLMs.
To overcome these challenges, we propose the Federated Prompt Tuning Paradigm for Multilingual Scenarios, which leverages parameter-efficient fine-tuning in a manner that preserves user privacy. We have designed a comprehensive set of experiments and introduced the concept of "language distance" to highlight the several strengths of this paradigm. Even under computational constraints, our method not only bolsters data efficiency but also facilitates mutual enhancements across languages, particularly benefiting low-resource ones. Compared to traditional local crosslingual transfer tuning methods, our approach achieves a 6.9\% higher accuracy, reduces the training parameters by over 99\%, and demonstrates stronger cross-lingual generalization. Such findings underscore the potential of our approach to promote social equality, ensure user privacy, and champion linguistic diversity. Wanru Zhao, Royson Lee, Xinchi Qiu, Yan Gao 0016, Hongxiang Fan, Nicholas D. Lane |
ICLR | 7 |
| 2024 | TinyTrain: Resource-Aware Task-Adaptive Sparse Training of DNNs at the Data-Scarce EdgeabstractOn-device training is essential for user personalisation and privacy. With the pervasiveness of IoT devices and microcontroller units (MCUs), this task becomes more challenging due to the constrained memory and compute resources, and the limited availability of labelled user data. Nonetheless, prior works neglect the data scarcity issue, require excessively long training time ($\textit{e.g.}$ a few hours), or induce substantial accuracy loss ($\geq$10%). In this paper, we propose TinyTrain, an on-device training approach that drastically reduces training time by selectively updating parts of the model and explicitly coping with data scarcity. TinyTrain introduces a task-adaptive sparse-update method that $\textit{dynamically}$ selects the layer/channel to update based on a multi-objective criterion that jointly captures user data, the memory, and the compute capabilities of the target device, leading to high accuracy on unseen tasks with reduced computation and memory footprint. TinyTrain outperforms vanilla fine-tuning of the entire network by 3.6-5.0% in accuracy, while reducing the backward-pass memory and computation cost by up to 1,098$\times$ and 7.68$\times$, respectively. Targeting broadly used real-world edge devices, TinyTrain achieves 9.5$\times$ faster and 3.5$\times$ more energy-efficient training over status-quo approaches, and 2.23$\times$ smaller memory footprint than SOTA methods, while remaining within the 1 MB memory envelope of MCU-grade platforms. Young D. Kwon, Rui Li 0052, Stylianos I. Venieris, Jagmohan Chauhan, Nicholas D. Lane, Cecilia Mascolo |
ICML | 5 |
| 2024 | Recurrent Early Exits for Federated Learning with Heterogeneous ClientsabstractFederated learning (FL) has enabled distributed learning of a model across multiple clients in a privacy-preserving manner. One of the main challenges of FL is to accommodate clients with varying hardware capacities; clients have differing compute and memory requirements. To tackle this challenge, recent state-of-the-art approaches leverage the use of early exits. Nonetheless, these approaches fall short of mitigating the challenges of joint learning multiple exit classifiers, often relying on hand-picked heuristic solutions for knowledge distillation among classifiers and/or utilizing additional layers for weaker classifiers. In this work, instead of utilizing multiple classifiers, we propose a recurrent early exit approach named ReeFL that fuses features from different sub-models into a single shared classifier. Specifically, we use a transformer-based early-exit module shared among sub-models to i) better exploit multi-layer feature representations for task-specific prediction and ii) modulate the feature representation of the backbone model for subsequent predictions. We additionally present a per-client self-distillation approach where the best sub-model is automatically selected as the teacher of the other sub-models at each client. Our experiments on standard image and speech classification benchmarks across various emerging federated fine-tuning baselines demonstrate ReeFL effectiveness over previous works. Royson Lee, Javier Fernández-Marqués, Shell Xu Hu, Da Li 0001, Stefanos Laskaridis, Lukasz Dudziak, Timothy M. Hospedales, Ferenc Huszar, Nicholas D. Lane |
ICML | 9 |
| 2024 | Towards Neural Architecture Search through Hierarchical Generative ModelingabstractNeural Architecture Search (NAS) aims to automate deep neural network design across various applications, while a good search space design is core to NAS performance. A too-narrow search space may fail to cover diverse task requirements, whereas a too-broad one can escalate computational expenses and reduce efficiency. %We propose automatically generating the search space to tailor it to specific task conditions, optimizing search costs and producing viable architectures. In this work, we aim to address this challenge by leaning on the recent advances in generative modelling -- we propose a novel method that can navigate through an extremely large, general-purpose initial search space efficiently by training a two-level generative model hierarchy. The first level uses Conditional Continuous Normalizing Flow (CCNF) for micro-cell design, while the second employs a transformer-based sequence generator to craft macro architectures aligned with task needs and architectural constraints. To ensure computational feasibility, we pretrain the generative models in a task-agnostic manner using a metric space of graph and zero-cost (ZC) similarities between architectures. We show our approach can achieve state-of-the-art performance among other low-cost NAS methods across different tasks on CIFAR-10/100, ImageNet and NAS-Bench-360. Lichuan Xiang, Lukasz Dudziak, Mohamed S. Abdelfattah, Abhinav Mehrotra, Nicholas D. Lane, Hongkai Wen 0001 |
ICML | 5 |
| 2024 | CLUES: Collaborative Private-domain High-quality Data Selection for LLMs via Training DynamicsabstractRecent research has highlighted the importance of data quality in scaling large language models (LLMs). However, automated data quality control faces unique challenges in collaborative settings where sharing is not allowed directly between data silos. To tackle this issue, this paper proposes a novel data quality control technique based on the notion of data influence on the training dynamics of LLMs, that high quality data are more likely to have similar training dynamics to the anchor dataset. We then leverage the influence of the training dynamics to select high-quality data from different private domains, with centralized model updates on the server side in a collaborative training fashion by either model merging or federated learning. As for the data quality indicator, we compute the per-sample gradients with respect to the private data and the anchor dataset, and use the trace of the accumulated inner products as a measurement of data quality. In addition, we develop a quality control evaluation tailored for collaborative settings with heterogeneous medical domain data. Experiments show that training on the high-quality data selected by our method can often outperform other data selection methods for collaborative fine-tuning of LLMs, across diverse private domain datasets, in medical, multilingual and financial settings. Wanru Zhao, Hongxiang Fan, Shell Xu Hu, Wangchunshu Zhou, Nicholas D. Lane |
NeurIPS | 5 |
| 2024 | Meta-Learned Kernel For Blind Super-Resolution Kernel EstimationabstractRecent image degradation estimation methods have enabled single-image super-resolution (SR) approaches to better upsample real-world images. Among these methods, explicit kernel estimation approaches have demonstrated unprecedented performance at handling unknown degradations. Nonetheless, a number of limitations constrain their efficacy when used by downstream SR models. Specifically, this family of methods yields i) excessive inference time due to long per-image adaptation times and ii) inferior image fidelity due to kernel mismatch. In this work, we introduce a learning-to-learn approach that meta-learns from the information contained in a distribution of images, thereby enabling significantly faster adaptation to new images with substantially improved performance in both kernel estimation and image fidelity. Specifically, we meta-train a kernelgenerating GAN, named MetaKernelGAN, on a range of tasks, such that when a new image is presented, the generator starts from an informed kernel estimate and the discriminator starts with a strong capability to distinguish between patch distributions. Compared with state-of-the-art methods, our experiments show that MetaKernelGAN better estimates the magnitude and covariance of the kernel, leading to state-of-the-art blind SR results within a similar computational regime when combined with a non-blind SR model. Through supervised learning of an unsupervised learner, our method maintains the generalizability of the unsupervised learner, improves the optimization stability of kernel estimation, and hence image adaptation, and leads to a faster inference with a speedup between 14.24 to 102.1× over existing methods.0 Royson Lee, Rui Li 0052, Stylianos I. Venieris, Timothy M. Hospedales, Ferenc Huszar, Nicholas D. Lane |
WACV | 6 |
| 2024 | NAWQ-SR: A Hybrid-Precision NPU Engine for Efficient On-Device Super-ResolutionabstractIn recent years, image and video delivery systems have begun integrating deep learning super-resolution (SR) approaches, leveraging their unprecedented visual enhancement capabilities while reducing reliance on networking conditions. Nevertheless, deploying these solutions on mobile devices still remains an active challenge as SR models are excessively demanding with respect to workload and memory footprint. Despite recent progress on on-device SR frameworks, existing systems either penalize visual quality, lead to excessive energy consumption or make inefficient use of the available resources. This work presents NAWQ-SR, a novel framework for the efficient on-device execution of SR models. Through a novel hybrid-precision quantization technique and a runtime neural image codec, NAWQ-SR exploits the multi-precision capabilities of modern mobile NPUs in order to minimize latency, while meeting user-specified quality constraints. Moreover, NAWQ-SR selectively adapts the arithmetic precision at run time to equip the SR DNN's layers with wider representational power, improving visual quality beyond what was previously possible on NPUs.Altogether, NAWQ-SR achieves an average speedup of 7.9×, 3× and 1.91× over the state-of-the-art on-device SR systems that use heterogeneous processors (MobiSR), CPU (SplitSR) and NPU (XLSR), respectively.Furthermore, NAWQ-SR delivers an average of 3.2× speedup and 0.39 dB higher PSNR over status-quo INT8 NPU designs, but most importantly mitigates the negative effects of quantization on visual quality, setting a new state-of-the-art in the attainable quality of NPU-based SR. Stylianos I. Venieris, Mário Almeida, Royson Lee, Nicholas D. Lane |
IEEE Trans. Mob. Comput. | 4 |
| 2023 | Zero-Cost Operation Scoring in Differentiable Architecture SearchabstractWe formalize and analyze a fundamental component of dif- ferentiable neural architecture search (NAS): local “opera- tion scoring” at each operation choice. We view existing operation scoring functions as inexact proxies for accuracy, and we find that they perform poorly when analyzed empir- ically on NAS benchmarks. From this perspective, we intro- duce a novel perturbation-based zero-cost operation scor- ing (Zero-Cost-PT) approach, which utilizes zero-cost prox- ies that were recently studied in multi-trial NAS but de- grade significantly on larger search spaces, typical for dif- ferentiable NAS. We conduct a thorough empirical evalu- ation on a number of NAS benchmarks and large search spaces, from NAS-Bench-201, NAS-Bench-1Shot1, NAS- Bench-Macro, to DARTS-like and MobileNet-like spaces, showing significant improvements in both search time and accuracy. On the ImageNet classification task on the DARTS search space, our approach improved accuracy compared to the best current training-free methods (TE-NAS) while be- ing over 10× faster (total searching time 25 minutes on a single GPU), and observed significantly better transferabil- ity on architectures searched on the CIFAR-10 dataset with an accuracy increase of 1.8 pp. Our code is available at: https://github.com/zerocostptnas/zerocost operation score. Lichuan Xiang, Lukasz Dudziak, Mohamed S. Abdelfattah, Thomas C. P. Chau, Nicholas D. Lane, Hongkai Wen 0001 |
AAAI | 5 |
| 2023 | L-DAWA: Layer-wise Divergence Aware Weight Aggregation in Federated Self-Supervised Visual Representation LearningabstractThe ubiquity of camera-enabled devices has led to large amounts of unlabeled image data being produced at the edge. The integration of self-supervised learning (SSL) and federated learning (FL) into one coherent system can potentially offer data privacy guarantees while also advancing the quality and robustness of the learned visual representations without needing to move data around. However, client bias and divergence during FL aggregation caused by data heterogeneity limits the performance of learned visual representations on downstream tasks. In this paper, we propose a new aggregation strategy termed Layer-wise Divergence Aware Weight Aggregation (L-DAWA) to mitigate the influence of client bias and divergence during FL aggregation. The proposed method aggregates weights at the layer-level according to the measure of angular divergence between the clients’ model and the global model. Extensive experiments with cross-silo and cross-device settings on CIFAR-10/100 and Tiny ImageNet datasets demonstrate that our methods are effective and obtain new SOTA performance on both contrastive and non-contrastive SSL approaches. Yasar Abbas Ur Rehman, Yan Gao 0016, Pedro Porto Buarque de Gusmão, Mina Alibeigi, Nicholas D. Lane |
ICCV | 6 |
| 2023 | Sparse-DySta: Sparsity-Aware Dynamic and Static Scheduling for Sparse Multi-DNN WorkloadsabstractRunning multiple deep neural networks (DNNs) in parallel has become an emerging workload in both edge devices, such as mobile phones where multiple tasks serve a single user for daily activities, and data centers, where various requests are raised from millions of users, as seen with large language models. To reduce the costly computational and memory requirements of these workloads, various efficient sparsification approaches have been introduced, resulting in widespread sparsity across different types of DNN models. In this context, there is an emerging need for scheduling sparse multi-DNN workloads, a problem that is largely unexplored in previous literature. This paper systematically analyses the use-cases of multiple sparse DNNs and investigates the opportunities for optimizations. Based on these findings, we propose Dysta, a novel bi-level dynamic and static scheduler that utilizes both static sparsity patterns and dynamic sparsity information for the sparse multi-DNN scheduling. Both static and dynamic components of Dysta are jointly designed at the software and hardware levels, respectively, to improve and refine the scheduling approach. To facilitate future progress in the study of this class of workloads, we construct a public benchmark that contains sparse multi-DNN workloads across different deployment scenarios, spanning from mobile phones and AR/VR wearables to data centers. A comprehensive evaluation on the sparse multi-DNN benchmark demonstrates that our proposed approach outperforms the state-of-the-art methods with up to 10% decrease in latency constraint violation rate and nearly 4 × reduction in average normalized turnaround time. Our artifacts and code are publicly available at: https://github.com/SamsungLabs/Sparse-Multi-DNN-Scheduling. Hongxiang Fan, Stylianos I. Venieris, Alexandros Kouris, Nicholas D. Lane |
MICRO | 4 |
| 2023 | FedL2P: Federated Learning to PersonalizeabstractFederated learning (FL) research has made progress in developing algorithms for distributed learning of global models, as well as algorithms for local personalization of those common models to the specifics of each client’s local data distribution. However, different FL problems may require different personalization strategies, and it may not even be possible to define an effective one-size-fits-all personalization strategy for all clients: Depending on how similar each client’s optimal predictor is to that of the global model, different personalization strategies may be preferred. In this paper, we consider the federated meta-learning problem of learning personalization strategies. Specifically, we consider meta-nets that induce the batch-norm and learning rate parameters for each client given local data statistics. By learning these meta-nets through FL, we allow the whole FL network to collaborate in learning a customized personalization strategy for each client. Empirical results show that this framework improves on a range of standard hand-crafted personalization baselines in both label and feature shift situations. Royson Lee, Minyoung Kim 0001, Da Li 0001, Xinchi Qiu, Timothy M. Hospedales, Ferenc Huszar, Nicholas D. Lane |
NeurIPS | 7 |
| 2023 | FedVal: Different good or different bad in federated learning
Viktor Valadi, Xinchi Qiu, Pedro Porto Buarque de Gusmão, Nicholas D. Lane, Mina Alibeigi |
USENIX Security Symposium | 4 |
| 2023 | Decentralized Training of 3D Lane Detection with Automatic Labeling Using HD MapsabstractTo have competent 3D lane detection for real-world driving, a massive amount of data from all over the world is needed, but data collection and manual annotation are costly and time-consuming. The diversity of data collected by developmental cars might still be limited compared to the data collected by a large fleet of customer cars. Federated learning enables training models on edge without transferring data out of devices. However, training supervised learning tasks at the edge is directly tied to having access to high-quality labels, which is limited at the edge.In this paper, we propose a fully automatic method to generate 3D lane labels at the edge using a pre-recorded HD map to enable the federated training of the 3D lane detection model. As a reference, a semi-automatic method is applied for creating a 3D-lane dataset used as ground truth. Our experimental results show that the model can achieve comparable performance when training on the same dataset in both a centralized and a decentralized manner. And the models trained on semi-automatic labeled datasets slightly outperform those trained on fully-automatically labeled datasets. This study shows that a well-performing 3D lane detection model can be trained in a supervised and fully decentralized manner, and most importantly, data privacy at the edge is guaranteed. Yadong Mao, Zhuqi Xiao, Che-Tsung Lin, Pedro Porto Buarque de Gusmão, Nicholas D. Lane, Christopher Zach, Mina Alibeigi |
VTC2023-Spring | 5 |
| 2023 | A First Look into the Carbon Footprint of Federated LearningabstractDespite impressive results, deep learning-based technologies also raise severe privacy and environmental concerns induced by the training procedure often conducted in data centers. In response, alternatives to centralized training such as Federated Learning (FL) have emerged. FL is now starting to be deployed at a global scale by companies that must adhere to new legal demands and policies originating from governments and social groups advocating for privacy protection. However, the potential environmental impact related to FL remains unclear and unexplored. This article offers the first-ever systematic study of the carbon footprint of FL. We propose a rigorous model to quantify the carbon footprint, hence facilitating the investigation of the relationship between FL design and carbon emissions. We also compare the carbon footprint of FL to traditional centralized learning. Our findings show that, depending on the configuration, FL can emit up to two orders of magnitude more carbon than centralized training. However, in certain settings, it can be comparable to centralized learning due to the reduced energy consumption of embedded devices. Finally, we highlight and connect the results to the future challenges and trends in FL to reduce its environmental impact, including algorithms efficiency, hardware capabilities, and stronger industry transparency. Xinchi Qiu, Titouan Parcollet, Javier Fernández-Marqués, Pedro Porto Buarque de Gusmão, Yan Gao 0016, Daniel J. Beutel, Taner Topal, Akhil Mathur, Nicholas D. Lane |
J. Mach. Learn. Res. | 9 |
| 2023 | Mitigating Memory Wall Effects in CNN Engines with On-the-Fly Weights GenerationabstractThe unprecedented accuracy of convolutional neural networks (CNNs) across a broad range of AI tasks has led to their widespread deployment in mobile and embedded settings. In a pursuit for high-performance and energy-efficient inference, significant research effort has been invested in the design of field-programmable gate array (FPGA)–based CNN accelerators. In this context, single computation engines constitute a popular design approach that enables the deployment of diverse models without the overhead of fabric reconfiguration. Nevertheless, this flexibility often comes with significantly degraded performance on memory-bound layers and resource underutilisation due to the suboptimal mapping of certain layers on the engine’s fixed configuration. In this work, we investigate the implications in terms of CNN engine design for a class of models that introduce a pre-convolution stage to decompress the weights at runtime. We refer to these approaches as on-the-fly . This article presents unzipFPGA, a novel CNN inference system that counteracts the limitations of existing CNN engines. The proposed framework comprises a novel CNN hardware architecture that introduces a weights generator module that enables the on-chip on-the-fly generation of weights, alleviating the negative impact of limited bandwidth on memory-bound layers. We further enhance unzipFPGA with an automated hardware-aware methodology that tailors the weights generation mechanism to the target CNN-device pair, leading to an improved accuracy–performance balance. Finally, we introduce an input selective processing element (PE) design that balances the load between PEs in suboptimally mapped layers. Quantitative evaluation shows that the proposed framework yields hardware designs that achieve an average of 2.57× performance efficiency gain over highly optimised GPU designs for the same power constraints and up to 3.94× higher performance density over a diverse range of state-of-the-art FPGA-based CNN accelerators. Stylianos I. Venieris, Javier Fernández-Marqués, Nicholas D. Lane |
ACM Trans. Design Autom. Electr. Syst. | 3 |
| 2022 | Multi-Exit Semantic Segmentation Networks
Alexandros Kouris, Stylianos I. Venieris, Stefanos Laskaridis, Nicholas D. Lane |
ECCV (21) | 4 |
| 2022 | Federated Self-supervised Learning for Video Understanding
Yasar Abbas Ur Rehman, Yan Gao 0016, Pedro Porto Buarque de Gusmão, Nicholas D. Lane |
ECCV (31) | 5 |
| 2022 | End-to-End Speech Recognition from Federated Acoustic ModelsabstractTraining Automatic Speech Recognition (ASR) models under federated learning (FL) settings has attracted a lot of attention recently. However, the FL scenarios often presented in the literature are artificial and fail to capture the complexity of real FL systems. In this paper, we construct a challenging and realistic ASR federated experimental setup consisting of clients with heterogeneous data distributions using the French and Italian sets of the CommonVoice dataset, a large heterogeneous dataset containing thousands of different speakers, acoustic environments and noises. We present the first empirical study on an attention-based sequence-to-sequence End-to-End (E2E) ASR model with three aggregation weighting strategies – standard FedAvg, loss-based aggregation and a novel word error rate (WER)-based aggregation, compared in two realistic FL scenarios: cross-silo with 10 clients and cross-device with 2K and 4K clients. This 4K cross-device ASR experiment is the largest ever performed. Our first-of-its-kind analysis on E2E ASR from heterogeneous and realistic federated acoustic models provides the foundations for future research and development of realistic FL ASR applications. Yan Gao 0016, Titouan Parcollet, Mohamed Salah Zaïem, Javier Fernández-Marqués, Pedro Porto Buarque de Gusmão, Daniel J. Beutel, Nicholas D. Lane |
ICASSP | 7 |
| 2022 | Prospect Pruning: Finding Trainable Weights at Initialization using Meta-Gradients
Milad Alizadeh, Shyam A. Tailor, Luisa M. Zintgraf, Joost van Amersfoort, Sebastian Farquhar, Nicholas D. Lane, Yarin Gal |
ICLR | 6 |
| 2022 | ZeroFL: Efficient On-Device Training for Federated Learning with Local Sparsity
Xinchi Qiu, Javier Fernández-Marqués, Pedro Porto Buarque de Gusmão, Yan Gao 0016, Titouan Parcollet, Nicholas D. Lane |
ICLR | 6 |
| 2022 | Conditioning Sequence-to-sequence Networks with Learned Activations
Alberto Gil C. P. Ramos, Abhinav Mehrotra, Nicholas D. Lane, Sourav Bhattacharya |
ICLR | 3 |
| 2022 | Do We Need Anisotropic Graph Neural Networks?
Shyam A. Tailor, Felix L. Opolka, Pietro Liò, Nicholas D. Lane |
ICLR | 4 |
| 2022 | Federated Self-supervised Speech Representations: Are We There Yet?
Yan Gao 0016, Javier Fernández-Marqués, Titouan Parcollet, Abhinav Mehrotra, Nicholas D. Lane |
INTERSPEECH | 5 |
| 2022 | Adaptable Butterfly Accelerator for Attention-based NNs via Hardware and Algorithm Co-designabstractAttention-based neural networks have become pervasive in many AI tasks. Despite their excellent algorithmic performance, the use of the attention mechanism and feedforward network (FFN) demands excessive computational and memory resources, which often compromises their hardware performance. Although various sparse variants have been introduced, most approaches only focus on mitigating the quadratic scaling of attention on the algorithm level, without explicitly considering the efficiency of mapping their methods on real hardware designs. Furthermore, most efforts only focus on either the attention mechanism or the FFNs but without jointly optimizing both parts, causing most of the current designs to lack scalability when dealing with different input lengths. This paper systematically considers the sparsity patterns in different variants from a hardware perspective. On the algorithmic level, we propose FABNet, a hardware-friendly variant that adopts a unified butterfly sparsity pattern to approximate both the attention mechanism and the FFNs. On the hardware level, a novel adaptable butterfly accelerator is proposed that can be configured at runtime via dedicated hardware control to accelerate different butterfly layers using a single unified hardware engine. On the Long-Range-Arena dataset, FABNet achieves the same accuracy as the vanilla Transformer while reducing the amount of computation by 10$\sim66\times$ and the number of parameters 2$\sim22\times$. By jointly optimizing the algorithm and hardware, our FPGA-based butterfly accelerator achieves 14.2$\sim23.2\times$ speedup over state-of-the-art accelerators normalized to the same computational budget. Compared with optimized CPU and GPU designs on Raspberry Pi 4 and Jetson Nano, our system is up to $273.8\times$ and $15.1\times$ faster under the same power budget Hongxiang Fan, Thomas C. P. Chau, Stylianos I. Venieris, Royson Lee, Alexandros Kouris, Wayne Luk, Nicholas D. Lane, Mohamed S. Abdelfattah |
MICRO | 7 |
| 2022 | Adaptable mobile vision systems through multi-exit neural networksabstractSemantic segmentation constitutes the backbone of many mobile vision systems, spanning from robot navigation to augmented reality and teleconferencing. Frequently operating under stringent latency constraints within the limited resource envelope of embedded/mobile devices, optimising for efficient execution becomes important. To this end, we propose a framework for converting state-of-the-art segmentation models to MESS networks: specially trained CNNs that employ parametrised early exits along their depth. Upon deployment, the predictions of these exits can be exploited either in a dynamic (input-adaptive) way, to save computation during inference on easier samples; or in a static (device-adaptive) setting, to accommodate deployment under varying device capabilities without the need of retraining. Designing and training such networks naively can hurt performance. Thus, we propose a two-staged training process that pushes semantically important features early in the network. We co-optimise the number, placement and architecture of the attached segmentation heads, along with the exit policy, to adapt to the deployment scenario and application-specific requirements. Optimising for speed, MESS networks deliver latency gains of up to 2.65× over state-of-the-art methods with no accuracy degradation. Accordingly, optimising for accuracy, we achieve an improvement of up to 5.33 pp, under the same computational budget. Alexandros Kouris, Stylianos I. Venieris, Stefanos Laskaridis, Nicholas D. Lane |
MobiSys | 4 |
| 2022 | BLOX: Macro Neural Architecture Search Benchmark and AlgorithmsabstractNeural architecture search (NAS) has been successfully used to design numerous high-performance neural networks. However, NAS is typically compute-intensive, so most existing approaches restrict the search to decide the operations and topological structure of a single block only, then the same block is stacked repeatedly to form an end-to-end model. Although such an approach reduces the size of search space, recent studies show that a macro search space, which allows blocks in a model to be different, can lead to better performance. To provide a systematic study of the performance of NAS algorithms on a macro search space, we release Blox – a benchmark that consists of 91k unique models trained on the CIFAR-100 dataset. The dataset also includes runtime measurements of all the models on a diverse set of hardware platforms. We perform extensive experiments to compare existing algorithms that are well studied on cell-based search spaces, with the emerging blockwise approaches that aim to make NAS scalable to much larger macro search spaces. The Blox benchmark and code are available at https://github.com/SamsungLabs/blox. Thomas C. P. Chau, Lukasz Dudziak, Hongkai Wen 0001, Nicholas D. Lane, Mohamed S. Abdelfattah |
NeurIPS | 4 |
| 2022 | Match to Win: Analysing Sequences Lengths for Efficient Self-Supervised Learning in Speech and AudioabstractSelf-supervised learning (SSL) has proven vital in speech and audio-related applications. The paradigm trains a general model on unlabeled data that can later be used to solve specific downstream tasks. This type of model is costly to train as it requires manipulating long input sequences that can only be handled by powerful centralised servers. Surprisingly, despite many attempts to increase training efficiency through model compression, the effects of truncating input sequence lengths to reduce computation have not been studied. In this paper, we provide the first empirical study of SSL pre-training for different specified sequence lengths and link this to various downstream tasks. We find that training on short sequences can dramatically reduce resource costs while retaining a satisfactory performance for all tasks. This simple one-line change would promote the migration of SSL training from data centres to user-end edge devices for more realistic and personalised applications. Yan Gao 0016, Javier Fernández-Marqués, Titouan Parcollet, Pedro Porto Buarque de Gusmão, Nicholas D. Lane |
SLT | 5 |
| 2022 | DynO: Dynamic Onloading of Deep Neural Networks from Cloud to DeviceabstractRecently, there has been an explosive growth of mobile and embedded applications using convolutional neural networks (CNNs). To alleviate their excessive computational demands, developers have traditionally resorted to cloud offloading, inducing high infrastructure costs and a strong dependence on networking conditions. On the other end, the emergence of powerful SoCs is gradually enabling on-device execution. Nonetheless, low- and mid-tier platforms still struggle to run state-of-the-art CNNs sufficiently. In this article, we present DynO, a distributed inference framework that combines the best of both worlds to address several challenges, such as device heterogeneity, varying bandwidth, and multi-objective requirements. Key components that enable this are its novel CNN-specific data packing method, which exploits the variability of precision needs in different parts of the CNN when onloading computation, and its novel scheduler, which jointly tunes the partition point and transferred data precision at runtime to adapt inference to its execution environment. Quantitative evaluation shows that DynO outperforms the current state of the art, improving throughput by over an order of magnitude over device-only execution and up to 7.9× over competing CNN offloading systems, with up to 60× less data transferred. Mário Almeida, Stefanos Laskaridis, Stylianos I. Venieris, Ilias Leontiadis, Nicholas D. Lane |
ACM Trans. Embed. Comput. Syst. | 5 |
| 2022 | Wireless Localization with Spatial-Temporal Robust FingerprintsabstractIndoor localization has gained increasing attention in the era of the Internet of Things. Among various technologies, WiFi fingerprint-based localization has become a mainstream solution. However, RSS fingerprints suffer from critical drawbacks of spatial ambiguity and temporal instability that root in multipath effects and environmental dynamics, which degrade the performance of these systems and therefore impede their wide deployment in the real world. Pioneering works overcome these limitations at the costs of ubiquity as they mostly resort to additional information or extra user constraints. In this article, we present the design and implementation of ViViPlus, an indoor localization system purely based on WiFi fingerprints, which jointly mitigates spatial ambiguity and temporal instability and derives reliable performance without impairing the ubiquity. The key idea is to embrace the spatial awareness of RSS values in a novel form of RSS Spatial Gradient (RSG) matrix for enhanced WiFi fingerprints. We devise techniques for the representation, construction, and localization of the proposed fingerprint form and integrate them all in a practical system. Extensive experiments across 7 months in different environments demonstrate that ViViPlus significantly improves the accuracy in localization scenarios by about 30% to 50% compared with the state-of-the-art approaches. Danyang Li 0005, Jingao Xu, Zheng Yang 0002, Chenshu Wu, Nicholas D. Lane |
ACM Trans. Sens. Networks | 6 |
| 2022 | FRuDA: Framework for Distributed Adversarial Domain AdaptationabstractBreakthroughs in unsupervised domain adaptation (uDA) can help in adapting models from a label-rich source domain to unlabeled target domains. Despite these advancements, there is a lack of research on how uDA algorithms, particularly those based on adversarial learning, can work in distributed settings. In real-world applications, target domains are often distributed across thousands of devices, and existing adversarial uDA algorithms – which are centralized in nature – cannot be applied in these settings. To solve this important problem, we introduce FRuDA: an end-to-end framework for distributed adversarial uDA. Through a careful analysis of the uDA literature, we identify the design goals for a distributed uDA system and propose two novel algorithms to increase adaptation accuracy and training efficiency of adversarial uDA in distributed settings. Our evaluation of FRuDA with five image and speech datasets show that it can boost target domain accuracy by up to 50% and improve the training efficiency of adversarial uDA by at least$11\times$. Shaoduo Gan, Akhil Mathur, Anton Isopoussu, Fahim Kawsar, Nadia Bianchi-Berthouze, Nicholas D. Lane |
IEEE Trans. Parallel Distributed Syst. | 6 |
| 2021 | Distilling Knowledge from Ensembles of Acoustic Models for Joint CTC-Attention End-to-End Speech RecognitionabstractKnowledge distillation has been widely used to compress existing deep learning models while preserving the performance on a wide range of applications. In the specific context of Automatic Speech Recognition (ASR), distillation from ensembles of acoustic models has recently shown promising results in increasing recognition performance. In this paper, we propose an extension of multi-teacher distillation methods to joint CTC-attention end-to-end ASR systems. We also introduce three novel distillation strategies. The core intuition behind them is to integrate the error rate metric to the teacher selection rather than solely focusing on the observed losses. In this way, we directly distill and optimize the student toward the relevant metric for speech recognition. We evaluate these strategies under a selection of training procedures on different datasets (TIMIT, Librispeech, Common Voice) and various languages (English, French, Italian). In particular, state-of-the-art error rates are reported on the Common Voice French, Italian and TIMIT datasets. Yan Gao 0016, Titouan Parcollet, Nicholas D. Lane |
ASRU | 3 |
| 2021 | Defensive Tensorization
Adrian Bulat, Jean Kossaifi, Sourav Bhattacharya, Yannis Panagakis, Timothy M. Hospedales, Georgios Tzimiropoulos, Nicholas D. Lane, Maja Pantic |
BMVC | 7 |
| 2021 | unzipFPGA: Enhancing FPGA-based CNN Engines with On-the-Fly Weights GenerationabstractSingle computation engines have become a popular design choice for FPGA-based convolutional neural networks (CNNs) enabling the deployment of diverse models without fabric reconfiguration. This flexibility, however, often comes with significantly reduced performance on memory-bound layers and resource underutilisation due to suboptimal mapping of certain layers on the engine's fixed configuration. In this work, we investigate the implications in terms of CNN engine design for a class of models that introduce a pre-convolution stage to decompress the weights at run time. We refer to these approaches as on-the-fly. To minimise the negative impact of limited bandwidth on memory-bound layers, we present a novel hardware component that enables the on-chip on-the-fly generation of weights. We further introduce an input selective processing element (PE) design that balances the load between PEs on suboptimally mapped layers. Finally, we present unzipFPGA, a framework to train on-the-fly models and traverse the design space to select the highest performing CNN engine configuration. Quantitative evaluation shows that unzipFPGA yields an average speedup of 2.14× and 71% over optimised status-quo and pruned CNN engines under constrained bandwidth and up to 3.69× higher performance density over the state-of-the-art FPGA-based CNN accelerators. Stylianos I. Venieris, Javier Fernández-Marqués, Nicholas D. Lane |
FCCM | 3 |
| 2021 | Zero-Cost Proxies for Lightweight NAS
Mohamed S. Abdelfattah, Abhinav Mehrotra, Lukasz Dudziak, Nicholas D. Lane |
ICLR | 4 |
| 2021 | NAS-Bench-ASR: Reproducible Neural Architecture Search for Speech Recognition
Abhinav Mehrotra, Alberto Gil C. P. Ramos, Sourav Bhattacharya, Lukasz Dudziak, Ravichander Vipperla, Thomas C. P. Chau, Mohamed S. Abdelfattah, Samin Ishtiaq, Nicholas D. Lane |
ICLR | 9 |
| 2021 | Degree-Quant: Quantization-Aware Training for Graph Neural Networks
Shyam A. Tailor, Javier Fernández-Marqués, Nicholas D. Lane |
ICLR | 3 |
| 2021 | Smart at what cost?: characterising mobile deep neural networks in the wildabstractWith smartphones' omnipresence in people's pockets, Machine Learning (ML) on mobile is gaining traction as devices become more powerful. With applications ranging from visual filters to voice assistants, intelligence on mobile comes in many forms and facets. However, Deep Neural Network (DNN) inference remains a compute intensive workload, with devices struggling to support intelligence at the cost of responsiveness. On the one hand, there is significant research on reducing model runtime requirements and supporting deployment on embedded devices. On the other hand, the strive to maximise the accuracy of a task is supported by deeper and wider neural networks, making mobile deployment of state-of-the-art DNNs a moving target. Mário Almeida, Stefanos Laskaridis, Abhinav Mehrotra, Lukasz Dudziak, Ilias Leontiadis, Nicholas D. Lane |
Internet Measurement Conference | 6 |
| 2021 | FjORD: Fair and Accurate Federated Learning under heterogeneous targets with Ordered DropoutabstractFederated Learning (FL) has been gaining significant traction across different ML tasks, ranging from vision to keyboard predictions. In large-scale deployments, client heterogeneity is a fact and constitutes a primary problem for fairness, training performance and accuracy. Although significant efforts have been made into tackling statistical data heterogeneity, the diversity in the processing capabilities and network bandwidth of clients, termed system heterogeneity, has remained largely unexplored. Current solutions either disregard a large portion of available devices or set a uniform limit on the model's capacity, restricted by the least capable participants.In this work, we introduce Ordered Dropout, a mechanism that achieves an ordered, nested representation of knowledge in Neural Networks and enables the extraction of lower footprint submodels without the need for retraining. We further show that for linear maps our Ordered Dropout is equivalent to SVD. We employ this technique, along with a self-distillation methodology, in the realm of FL in a framework called FjORD. FjORD alleviates the problem of client system heterogeneity by tailoring the model width to the client's capabilities. Extensive evaluation on both CNNs and RNNs across diverse modalities shows that FjORD consistently leads to significant performance gains over state-of-the-art baselines while maintaining its nested structure. Samuel Horváth, Stefanos Laskaridis, Mário Almeida, Ilias Leontiadis, Stylianos I. Venieris, Nicholas D. Lane |
NeurIPS | 6 |
| 2021 | Chronic Pain Protective Behavior Detection with Deep LearningabstractIn chronic pain rehabilitation, physiotherapists adapt physical activity to patients’ performance based on their expression of protective behavior, gradually exposing them to feared but harmless and essential everyday activities. As rehabilitation moves outside the clinic, technology should automatically detect such behavior to provide similar support. Previous works have shown the feasibility of automatic protective behavior detection (PBD) within a specific activity. In this article, we investigate the use of deep learning for PBD across activity types, using wearable motion capture and surface electromyography data collected from healthy participants and people with chronic pain. We approach the problem by continuously detecting protective behavior within an activity rather than estimating its overall presence. The best performance reaches mean F1 score of 0.82 with leave-one-subject-out cross validation. When protective behavior is modeled per activity type, performance achieves a mean F1 score of 0.77 for bend-down, 0.81 for one-leg-stand, 0.72 for sit-to-stand, 0.83 for stand-to-sit, and 0.67 for reach-forward. This performance reaches excellent level of agreement with the average experts’ rating performance suggesting potential for personalized chronic pain management at home. We analyze various parameters characterizing our approach to understand how the results could generalize to other PBD datasets and different levels of ground truth granularity. Temitayo A. Olugbade, Akhil Mathur, Amanda C. de C. Williams, Nicholas D. Lane, Nadia Bianchi-Berthouze |
ACM Trans. Comput. Heal. | 5 |
| 2020 | Best of Both Worlds: AutoML Codesign of a CNN and its Hardware AcceleratorabstractNeural architecture search (NAS) has been very successful at outperforming human-designed convolutional neural networks (CNN) in accuracy, and when hardware information is present, latency as well. However, NAS-designed CNNs typically have a complicated topology, therefore, it may be difficult to design a custom hardware (HW) accelerator for such CNNs. We automate HW-CNN codesign using NAS by including parameters from both the CNN model and the HW accelerator, and we jointly search for the best model-accelerator pair that boosts accuracy and efficiency. We call this Codesign-NAS. In this paper we focus on defining the Codesign-NAS multiobjective optimization problem, demonstrating its effectiveness, and exploring different ways of navigating the codesign search space. For CIFAR-10 image classification, we enumerate close to 4 billion model-accelerator pairs, and find the Pareto frontier within that large search space. This allows us to evaluate three different reinforcement-learning-based search strategies. Finally, compared to ResNet on its most optimal HW accelerator from within our HW design space, we improve on CIFAR-100 classification accuracy by 1.3% while simultaneously increasing performance/area by 41% in just ~1000 GPU-hours of running Codesign-NAS. Mohamed S. Abdelfattah, Lukasz Dudziak, Thomas C. P. Chau, Royson Lee, Hyeji Kim, Nicholas D. Lane |
DAC | 6 |
| 2020 | Journey Towards Tiny Perceptual Super-Resolution
Royson Lee, Lukasz Dudziak, Mohamed S. Abdelfattah, Stylianos I. Venieris, Hyeji Kim, Hongkai Wen 0001, Nicholas D. Lane |
ECCV (26) | 7 |
| 2020 | EMOPAIN Challenge 2020: Multimodal Pain Evaluation from Facial and Bodily ExpressionsabstractThe EmoPain 2020 Challenge is the first international competition aimed at creating a uniform platform for the comparison of multi-modal machine learning and multimedia processing methods of chronic pain assessment from human expressive behaviour, and also the identification of pain-related behaviours. The objective of the challenge is to promote research in the development of assistive technologies that help improve the quality of life for people with chronic pain via real-time monitoring and feedback to help manage their condition and remain physically active. The challenge also aims to encourage the use of the relatively underutilised, albeit vital bodily expression signals for automatic pain and pain-related emotion recognition. This paper presents a description of the challenge, competition guidelines, bench-marking dataset, and the baseline systems' architecture and performance on the Challenge's three sub-tasks: pain estimation from facial expressions, pain recognition from multimodal movement, and protective movement behaviour detection. Joy Egede, Siyang Song, Temitayo A. Olugbade, Amanda C. de C. Williams, Hongying Meng, M. S. Hane Aung, Nicholas D. Lane, Michel F. Valstar, Nadia Bianchi-Berthouze |
FG | 8 |
| 2020 | Codesign-NAS: Automatic FPGA/CNN Codesign Using Neural Architecture SearchabstractField-programmable gate arrays (FPGAs) have become a popular compute platform for convolutional neural network (CNN) inference; however, the design of a CNN model and its FPGA accelerator has been inherently sequential. A CNN is first prototyped with no-or-little hardware awareness to attain high accuracy; subsequently, an FPGA accelerator is tuned to that specific CNN to maximize its efficiency. Instead, we formulate a neural architecture search (NAS) optimization problem that contains parameters from both the CNN model and the FPGA accelerator, and we jointly search for the best CNN model-accelerator pair that boosts accuracy and efficiency -we call this Codesign-NAS. In this paper we focus on defining the Codesign-NAS multiobjective optimization problem, demonstrating its effectiveness, and exploring different ways of navigating the codesign search space. For Cifar-10 image classification, we enumerate close to 4 billion model-accelerator pairs, and find the Pareto frontier within that large search space. Next we propose accelerator innovations that improve the entire Pareto frontier. Finally, we compare to ResNet on a highly-tuned accelerator, and show that using codesign, we can improve on Cifar-100 classification accuracy by 1.8% while simultaneously increasing performance/area by 41% in just 1000 GPU-hours of running Codesign-NAS, thus demonstrating that our automated codesign approach is superior to sequential design of a CNN model and accelerator. Mohamed S. Abdelfattah, Lukasz Dudziak, Thomas C. P. Chau, Royson Lee, Hyeji Kim, Nicholas D. Lane |
FPGA | 6 |
| 2020 | Libri-Adapt: a New Speech Dataset for Unsupervised Domain AdaptationabstractThis paper introduces a new dataset, Libri-Adapt, to support unsupervised domain adaptation research on speech recognition models. Built on top of the LibriSpeech corpus, Libri-Adapt contains 7200 hours of English speech recorded on mobile and embedded-scale microphones, and spans 72 different domains that are representative of the challenging practical scenarios encountered by ASR models. More specifically, Libri-Adapt facilitates the study of domain shifts in ASR models caused by a) different acoustic environments, b) variations in speaker accents, c) previously unexplored factors such as heterogeneity in the hardware and platform software of the microphones, and d) a combination of the aforementioned three shifts. We also provide a number of baseline results quantifying the impact of these domain shifts on the Mozilla DeepSpeech2 ASR model. Akhil Mathur, Fahim Kawsar, Nadia Bianchi-Berthouze, Nicholas D. Lane |
ICASSP | 4 |
| 2020 | HAPI: Hardware-Aware Progressive InferenceabstractConvolutional neural networks (CNNs) have recently become the state-of-the-art in a diversity of AI tasks. Despite their popularity, CNN inference still comes at a high computational cost. A growing body of work aims to alleviate this by exploiting the difference in the classification difficulty among samples and early-exiting at different stages of the network. Nevertheless, existing studies on early exiting have primarily focused on the training scheme, without considering the use-case requirements or the deployment platform. This work presents HAPI, a novel methodology for generating high-performance early-exit networks by co-optimising the placement of intermediate exits together with the early-exit strategy at inference time. Furthermore, we propose an efficient design space exploration algorithm which enables the faster traversal of a large number of alternative architectures and generates the highest-performing design, tailored to the use-case requirements and target hardware. Quantitative evaluation shows that our system consistently outperforms alternative search mechanisms and state-of-the-art early-exit schemes across various latency budgets. Moreover, it pushes further the performance of highly optimised hand-crafted early-exit CNNs, delivering up to 5.11× speedup over lightweight models on imposed latency-driven SLAs for embedded devices. Stefanos Laskaridis, Stylianos I. Venieris, Hyeji Kim, Nicholas D. Lane |
ICCAD | 4 |
| 2020 | Unsupervised Domain Adaptation Under Label Space Mismatch for Speech ClassificationabstractUnsupervised domain adaptation using adversarial learning has shown promise in adapting speech models from a labeled source domain to an unlabeled target domain. However, prior works make a strong assumption that the label spaces of source and target domains are identical, which can be easily violated in real-world conditions. We present AMLS, an end-to-end architecture that performs Adaptation under Mismatched Label Spaces using two weighting schemes to separate shared and private classes in each domain. An evaluation on three speech adaptation tasks, namely gender, microphone, and emotion adaptation, shows that AMLS provides significant accuracy gains over baselines used in speech and vision adaptation tasks. Our contribution paves the way for applying UDA to speech models in unconstrained settings with no assumptions on the source and target label spaces. Akhil Mathur, Nadia Bianchi-Berthouze, Nicholas D. Lane |
INTERSPEECH | 3 |
| 2020 | Iterative Compression of End-to-End ASR Model Using AutoMLabstractIncreasing demand for on-device Automatic Speech Recognition (ASR) systems has resulted in renewed interests in developing automatic model compression techniques. Past research have shown that AutoML-based Low Rank Factorization (LRF) technique, when applied to an end-to-end Encoder-Attention-Decoder style ASR model, can achieve a speedup of up to 3.7x, outperforming laborious manual rank-selection approaches. However, we show that current AutoML-based search techniques only work up to a certain compression level, beyond which they fail to produce compressed models with acceptable word error rates (WER). In this work, we propose an iterative AutoML-based LRF approach that achieves over 5x compression without degrading the WER, thereby advancing the state-of-the-art in ASR compression. Abhinav Mehrotra, Lukasz Dudziak, Jinsu Yeo, Young-Yoon Lee, Ravichander Vipperla, Mohamed S. Abdelfattah, Sourav Bhattacharya, Samin Ishtiaq, Alberto Gil C. P. Ramos, Nicholas D. Lane |
INTERSPEECH | 12 |
| 2020 | FusionRNN: Shared Neural Parameters for Multi-Channel Distant Speech Recognition
Titouan Parcollet, Xinchi Qiu, Nicholas D. Lane |
INTERSPEECH | 3 |
| 2020 | Quaternion Neural Networks for Multi-Channel Distant Speech RecognitionabstractInternational audience Xinchi Qiu, Titouan Parcollet, Mirco Ravanelli, Nicholas D. Lane, Mohamed Morchid |
INTERSPEECH | 4 |
| 2020 | Bunched LPCNet: Vocoder for Low-Cost Neural Text-To-Speech SystemsabstractLPCNet is an efficient vocoder that combines linear prediction and deep neural network modules to keep the computational complexity low. In this work, we present two techniques to further reduce it's complexity, aiming for a low-cost LPCNet vocoder-based neural Text-to-Speech (TTS) System. These techniques are: 1) Sample-bunching, which allows LPCNet to generate more than one audio sample per inference; and 2) Bit-bunching, which reduces the computations in the final layer of LPCNet. With the proposed bunching techniques, LPCNet, in conjunction with a Deep Convolutional TTS (DCTTS) acoustic model, shows a 2.19x improvement over the baseline run-time when running on a mobile device, with a less than 0.1 decrease in TTS mean opinion score (MOS). Ravichander Vipperla, Kihyun Choo, Samin Ishtiaq, Kyoungbo Min, Sourav Bhattacharya, Abhinav Mehrotra, Alberto Gil C. P. Ramos, Nicholas D. Lane |
INTERSPEECH | 9 |
| 2020 | SPINN: synergistic progressive inference of neural networks over device and cloudabstractDespite the soaring use of convolutional neural networks (CNNs) in mobile applications, uniformly sustaining high-performance inference on mobile has been elusive due to the excessive computational demands of modern CNNs and the increasing diversity of deployed devices. A popular alternative comprises offloading CNN processing to powerful cloud-based servers. Nevertheless, by relying on the cloud to produce outputs, emerging mission-critical and high-mobility applications, such as drone obstacle avoidance or interactive applications, can suffer from the dynamic connectivity conditions and the uncertain availability of the cloud. In this paper, we propose SPINN, a distributed inference system that employs synergistic device-cloud computation together with a progressive inference method to deliver fast and robust CNN inference across diverse settings. The proposed system introduces a novel scheduler that co-optimises the early-exit policy and the CNN splitting at run time, in order to adapt to dynamic conditions and meet user-defined service-level requirements. Quantitative evaluation illustrates that SPINN outperforms its state-of-the-art collaborative inference counterparts by up to 2× in achieved throughput under varying network conditions, reduces the server cost by up to 6.8× and improves accuracy by 20.7% under latency constraints, while providing robust operation under uncertain connectivity conditions and significant energy savings compared to cloud-centric execution. Stefanos Laskaridis, Stylianos I. Venieris, Mário Almeida, Ilias Leontiadis, Nicholas D. Lane |
MobiCom | 5 |
| 2020 | BRP-NAS: Prediction-based NAS using GCNsabstractNeural architecture search (NAS) enables researchers to automatically explore broad design spaces in order to improve efficiency of neural networks. This efficiency is especially important in the case of on-device deployment, where improvements in accuracy should be balanced out with computational demands of a model. In practice, performance metrics of model are computationally expensive to obtain. Previous work uses a proxy (e.g., number of operations) or a layer-wise measurement of neural network layers to estimate end-to-end hardware performance but the imprecise prediction diminishes the quality of NAS. To address this problem, we propose BRP-NAS, an efficient hardware-aware NAS enabled by an accurate performance predictor-based on graph convolutional network (GCN). What is more, we investigate prediction quality on different metrics and show that sample efficiency of the predictor-based NAS can be improved by considering binary relations of models and an iterative data selection strategy. We show that our proposed method outperforms all prior methods on NAS-Bench-101, NAS-Bench-201 and DARTS. Finally, to raise awareness of the fact that accurate latency estimation is not a trivial task, we release LatBench -- a latency dataset of NAS-Bench-201 models running on a broad range of devices. Lukasz Dudziak, Thomas C. P. Chau, Mohamed S. Abdelfattah, Royson Lee, Hyeji Kim, Nicholas D. Lane |
NeurIPS | 6 |
| 2020 | A Hybrid Deep Learning Architecture for Privacy-Preserving Mobile AnalyticsabstractInternet-of-Things (IoT) devices and applications are being deployed in our homes and workplaces. These devices often rely on continuous data collection to feed machine learning models. However, this approach introduces several privacy and efficiency challenges, as the service operator can perform unwanted inferences on the available data. Recently, advances in edge processing have paved the way for more efficient, and private, data processing at the source for simple tasks and lighter models, though they remain a challenge for larger and more complicated models. In this article, we present a hybrid approach for breaking down large, complex deep neural networks for cooperative, and privacy-preserving analytics. To this end, instead of performing the whole operation on the cloud, we let an IoT device to run the initial layers of the neural network, and then send the output to the cloud to feed the remaining layers and produce the final result. In order to ensure that the user's device contains no extra information except what is necessary for the main task and preventing any secondary inference on the data, we introduce Siamese fine-tuning. We evaluate the privacy benefits of this approach based on the information exposed to the cloud service. We also assess the local inference cost of different layers on a modern handset. Our evaluations show that by using Siamese fine-tuning and at a small processing cost, we can greatly reduce the level of unnecessary, potentially sensitive information in the personal data, thus achieving the desired tradeoff between utility, privacy, and performance. Seyed Ali Ossia, Ali Shahin Shamsabadi, Sina Sajadmanesh, Ali Taheri, Kleomenis Katevas, Hamid R. Rabiee 0001, Nicholas D. Lane, Hamed Haddadi 0001 |
IEEE Internet Things J. | 7 |
| 2019 | Recurrent network based automatic detection of chronic pain protective behavior using MoCap and sEMG dataabstractIn chronic pain physical rehabilitation, physiotherapists adapt exercise sessions according to the movement behavior of patients. As rehabilitation moves beyond clinical sessions, technology is needed to similarly assess movement behaviors and provide such personalized support. In this paper, as a first step, we investigate automatic detection of protective behavior (movement behavior due to pain-related fear or pain) based on wearable motion capture and electromyography sensor data. We investigate two recurrent networks (RNN) referred to as stacked-LSTM and dual-stream LSTM, which we compare with related deep learning (DL) architectures. We further explore data augmentation techniques and additionally analyze the impact of segmentation window lengths on detection performance. The leading performance of 0.815 mean F1 score achieved by stacked-LSTM provides important grounding for the development of wearable technology to support chronic pain physical rehabilitation during daily activities. Temitayo A. Olugbade, Akhil Mathur, Amanda C. de C. Williams, Nicholas D. Lane, Nadia Bianchi-Berthouze |
UbiComp | 5 |
| 2019 | An Empirical study of Binary Neural Networks' Optimisation
Milad Alizadeh, Javier Fernández-Marqués, Nicholas D. Lane, Yarin Gal |
ICLR (Poster) | 3 |
| 2019 | FlexAdapt: Flexible Cycle-Consistent Adversarial Domain AdaptationabstractUnsupervised domain adaptation is emerging as a powerful technique to improve the generalizability of deep learning models to new image domains without using any labeled data in the target domain. In the literature, solutions which perform cross-domain feature-matching (e.g., ADDA), pixel-matching (CycleGAN), and combination of the two (e.g., CyCADA) have been proposed for unsupervised domain adaptation. Many of these approaches make a strong assumption that the source and target label spaces are the same, however in the real-world, this assumption does not hold true. In this paper, we propose a novel solution, FlexAdapt, which extends the state-of-the-art unsupervised domain adaptation approach of CyCADA to scenarios where the label spaces in source and target domains are only partially overlapped. Our solution beats a number of state-of-the-art baseline approaches by as much as 29% in some scenarios, and represent a way forward for applying domain adaptation techniques in the real world. Akhil Mathur, Anton Isopoussu, Fahim Kawsar, Nadia Bianchi-Berthouze, Nicholas D. Lane |
ICMLA | 5 |
| 2019 | ShrinkML: End-to-End ASR Model Compression Using Reinforcement LearningabstractEnd-to-end automatic speech recognition (ASR) models are increasingly large and complex to achieve the best possible accuracy. In this paper, we build an AutoML system that uses reinforcement learning (RL) to optimize the per-layer compression ratios when applied to a state-of-the-art attention based end-to-end ASR model composed of several LSTM layers. We use singular value decomposition (SVD) low-rank matrix factorization as the compression method. For our RL-based AutoML system, we focus on practical considerations such as the choice of the reward/punishment functions, the formation of an effective search space, and the creation of a representative but small data set for quick evaluation between search steps. Finally, we present accuracy results on LibriSpeech of the model compressed by our AutoML system, and we compare it to manually-compressed models. Our results show that in the absence of retraining our RL-based search is an effective and practical method to compress a production-grade ASR system. When retraining is possible, we show that our AutoML system can select better highly-compressed seed models compared to manually hand-crafted rank selection, thus allowing for more compression than previously possible. Lukasz Dudziak, Mohamed S. Abdelfattah, Ravichander Vipperla, Stefanos Laskaridis, Nicholas D. Lane |
INTERSPEECH | 5 |
| 2019 | Mic2Mic: using cycle-consistent generative adversarial networks to overcome microphone variability in speech systemsabstractMobile and embedded devices are increasingly using microphones and audio-based computational models to infer user context. A major challenge in building systems that combine audio models with commodity microphones is to guarantee their accuracy and robustness in the real-world. Besides many environmental dynamics, a primary factor that impacts the robustness of audio models is microphone variability. In this work, we propose Mic2Mic - a machine-learned system component - which resides in the inference pipeline of audio models and at real-time reduces the variability in audio data caused by microphone-specific factors. Two key considerations for the design of Mic2Mic were: a) to decouple the problem of microphone variability from the audio task, and b) put minimal burden on end-users to provide training data. With these in mind, we apply the principles of cycle-consistent generative adversarial networks (CycleGANs) to learn Mic2Mic using unlabeled and unpaired data collected from different microphones. Our experiments show that Mic2Mic can recover between 66% to 89% of the accuracy lost due to microphone variability for two common audio tasks. Akhil Mathur, Anton Isopoussu, Fahim Kawsar, Nadia Bianchi-Berthouze, Nicholas D. Lane |
IPSN | 5 |
| 2019 | MobiSR: Efficient On-Device Super-Resolution through Heterogeneous Mobile ProcessorsabstractIn recent years, convolutional networks have demonstrated unprecedented performance in the image restoration task of super-resolution (SR). SR entails the upscaling of a single low-resolution image in order to meet application-specific image quality demands and plays a key role in mobile devices. To comply with privacy regulations and reduce the overhead of cloud computing, executing SR models locally on-device constitutes a key alternative approach. Nevertheless, the excessive compute and memory requirements of SR workloads pose a challenge in mapping SR networks on resource-constrained mobile platforms. This work presents MobiSR, a novel framework for performing efficient super-resolution on-device. Given a target mobile platform, the proposed framework considers popular model compression techniques and traverses the design space to reach the highest performing trade-off between image quality and processing speed. At run time, a novel scheduler dispatches incoming image patches to the appropriate model-engine pair based on the patch's estimated upscaling difficulty in order to meet the required image quality with minimum processing latency. Quantitative evaluation shows that the proposed framework yields on-device SR designs that achieve an average speedup of 2.13x over highly-optimized parallel difficulty-unaware mappings and 4.79x over highly-optimized single compute engine implementations. Royson Lee, Stylianos I. Venieris, Lukasz Dudziak, Sourav Bhattacharya, Nicholas D. Lane |
MobiCom | 5 |
| 2019 | Poster: MobiSR - Efficient On-Device Super-Resolution through Heterogeneous Mobile ProcessorsabstractIn recent years, convolutional networks have demonstrated unprecedented performance in the image restoration task of super-resolution (SR). SR entails the upscaling of a single low-resolution image in order to meet application-specific image quality demands and plays a key role in mobile devices. To comply with privacy regulations and reduce the overhead of cloud computing, executing SR models locally on-device constitutes a key alternative approach. Nevertheless, the excessive compute and memory requirements of SR workloads pose a challenge in mapping SR networks on resource-constrained mobile platforms. This work presents MobiSR, a novel framework for performing efficient super-resolution on-device. Given a target mobile platform, the proposed framework considers popular model compression techniques and traverses the design space to reach the highest performing trade-off between image quality and processing speed. At run time, a novel scheduler dispatches incoming image patches to the appropriate model-engine pair based on the patch's estimated upscaling difficulty in order to meet the required image quality with minimum processing latency. Quantitative evaluation shows that the proposed framework yields on-device SR designs that achieve an average speedup of 2.13x over highly-optimized parallel difficulty-unaware mappings and 4.79x over highly-optimized single compute engine implementations. Royson Lee, Stylianos I. Venieris, Lukasz Dudziak, Sourav Bhattacharya, Nicholas D. Lane |
MobiCom | 5 |
| 2018 | Deterministic Binary Filters for Convolutional Neural NetworksabstractWe propose Deterministic Binary Filters, an approach to Convolutional Neural Networks that learns weighting coefficients of predefined orthogonal binary basis instead of the conventional approach of learning directly the convolutional filters. This approach results in model architectures with significantly fewer parameters (4x to 16x) and smaller model sizes (32x due to the use of binary rather than floating point precision). We show our deterministic filter design can be integrated into well-known network architectures (such as ResNet and SqueezeNet) with as little as 2% loss of accuracy (under datasets like CIFAR-10). Under ImageNet, they result in 3x model size reduction compared to sub-megabyte binary networks while reaching comparable accuracy levels. Vincent W. S. Tseng, Sourav Bhattacharya, Javier Fernández-Marqués, Milad Alizadeh, Catherine Tong, Nicholas D. Lane |
IJCAI | 6 |
| 2018 | Using deep data augmentation training to address software and hardware heterogeneities in wearable and smartphone sensing devicesabstractA small variation in mobile hardware and software can potentially cause a significant heterogeneity or variation in the sensor data each device collects. For example, the microphone and accelerometer sensors on different devices can respond very differently to the same audio or motion phenomena. Other factors, like the instantaneous computational load on a smartphone, can cause key behavior like sensor sampling rates to fluctuate, further polluting the data. When sensing devices are deployed in unconstrained and real-world conditions, examples of sharply lower classification accuracy are observed due to what is collectively known as the sensing system heterogeneity. In this work, we take an unconventional approach and argue against solving individual forms of heterogeneity, e.g., improving OS behavior, or the quality/uniformity of components. Instead, we propose and build classifiers that themselves are more tolerant of these variations by leveraging deep learning and a data-augmented training process. Neither augmentation nor deep learning has previously been attempted to cope with sensor heterogeneity. We systematically investigate how these two machine learning methodologies can be adapted to solve such problems, and identify when and where they are able to be successful. We find that our proposed approach is able to reduce classifier errors on an average by 9% and 17% for a range of inertial-and audio-based mobile classification tasks. Akhil Mathur, Sourav Bhattacharya, Petar Velickovic, Leonid Joffe, Nicholas D. Lane, Fahim Kawsar, Pietro Liò |
IPSN | 6 |
| 2018 | Embracing Spatial Awareness for Reliable WiFi-Based Indoor Location SystemsabstractIndoor localization gains increasingly attentions in the era of Internet of Things. Among various technologies, WiFi-based systems that leverage Received Signal Strengths (RSSs) as location fingerprints become the mainstream solutions. However, RSS fingerprints suffer from critical drawbacks of spatial ambiguity and temporal instability that root in multipath effects and environmental dynamics, which degrade the performance of these systems and therefore impede their wide deployment in real world. Pioneering works overcome these limitations at the costs of ubiquity as they mostly resort to additional information or extra user constraints. In this paper, we present the design and implementation of MatLoc, an indoor localization system purely based on WiFi fingerprints, which jointly mitigates spatial ambiguity and temporal instability and derives reliable performance without impairing the ubiquity. The key idea is to embrace the spatial awareness of RSS values in a novel form of RSS Spatial Gradient (RSG) matrix for enhanced WiFi fingerprints. We devise techniques for the representation, construction, and comparison of the proposed fingerprint form, and integrate them all in a practical system, which follows the classical fingerprinting framework and requires no more inputs than any previous RSS fingerprint based systems. Extensive experiments in different environments demonstrate that MatLoc significantly improves the accuracy in both localization and tracking scenarios by about 30% to 50% compared with five state-of-the-art approaches. Jingao Xu, Zheng Yang 0002, Hengjie Chen, Yunhao Liu 0001, Xiancun Zhou, Nicholas D. Lane |
MASS | 7 |
| 2018 | Using Pre-trained Full-Precision Models to Speed Up Training Binary Networks For Mobile DevicesabstractBinary Neural Networks (BNNs) are well-suited for deploying Deep Neural Networks (DNNs) to small embedded devices but state-of-the-art BNNs need to be trained from scratch for a long time. We show how weights from a pre-trained full-precision model can be used to speed-up training of binary networks. We show that for CIFAR-10, accuracies within 1% of the full-precision model can be achieved in just 5 epochs. Milad Alizadeh, Nicholas D. Lane |
MobiSys | 2 |
| 2018 | Deterministic binary filters for keyword spotting applicationsabstractWe present a binary architecture with 60% fewer parameters and 50% fewer operations during inference compared to the current state of the art for keyword spotting (KWS) applications at the cost of 3.4% accuracy. We construct convolutional filters on-the-fly using orthogonal binary codes and results in a compact architecture that would fit in devices with less than 30kB of memory. Javier Fernández-Marqués, Vincent W. S. Tseng, Sourav Bhattacharya, Nicholas D. Lane |
MobiSys | 4 |
| 2018 | Inference of Big-Five Personality Using Large-scale Networked Mobile and Appliance DataabstractWe present the first large-scale (9270-user) study of data from both mobile and networked appliances for Big-Five personality inference. We correlate aggregated behavioral and physical health features with personalities, and perform binary classification using SVM and Decision Tree. We find that it is possible to infer each Big-Five personality at accuracies of 75% from this dataset despite its size and complexity (mix of mobile and appliance) as prior methods offer similar accuracy levels. We would like to achieve better accuracies and this study is a first step towards seeing how to model such data. Catherine Tong, Gabriella M. Harari, Angela Chieh, Otmane Bellahsen, Matthieu Vegreville, Eva Roitmann, Nicholas D. Lane |
MobiSys | 7 |
| 2017 | Density-aware compressive crowdsensingabstractCrowdsensing systems collect large-scale sensor data from mobile devices to provide a wide-area view of phenomena including traffic, noise and air pollution. Because such data often exhibits sparse structure, it is natural to apply compressive sensing (CS) for data sampling and recovery. However in practice, crowd participants are often distributed highly unevenly across the sensing area, and thus the numbers of observations collected over different areas may vary wildly - an issue we call density disparity. Density disparity leads to inaccuracy in low density areas, and potentially undermines the recovery performance if conventional compressive sensing is applied directly, which equally treats data from areas of different density. Xiaohong Hao, Nicholas D. Lane, Xin Liu 0002, Thomas Moscibroda |
IPSN | 3 |
| 2017 | Accelerating Mobile Audio Sensing Algorithms through On-Chip GPU OffloadingabstractGPUs have recently enjoyed increased popularity as general purpose software accelerators in multiple application domains including computer vision and natural language processing. However, there has been little exploration into the performance and energy trade-offs mobile GPUs can deliver for the increasingly popular workload of deep-inference audio sensing tasks, such as, spoken keyword spotting in energy-constrained smartphones and wearables. In this paper, we study these trade-offs and introduce an optimization engine that leverages a series of structural and memory access optimization techniques that allow audio algorithm performance to be automatically tuned as a function of GPU device specifications and model semantics. We find that parameter optimized audio routines obtain inferences an order of magnitude faster than sequential CPU implementations, and up to 6.5x times faster than cloud offloading with good connectivity, while critically consuming 3-4x less energy than the CPU. Under our optimized GPU, conventional wisdom about how to use the cloud and low power chips is broken. Unless the network has a throughput of at least 20Mbps (and a RTT of 25 ms or less), with only about 10 to 20 seconds of buffering audio data for batched execution, the optimized GPU audio sensing apps begin to consume less energy than cloud offloading. Under such conditions we find the optimized GPU can provide energy benefits comparable to low-power reference DSP implementations with some preliminary level of optimization; in addition to the GPU always winning with lower latency. Petko Georgiev, Nicholas D. Lane, Cecilia Mascolo, David Chu |
MobiSys | 2 |
| 2017 | DeepEye: Resource Efficient Local Execution of Multiple Deep Vision Models using Wearable Commodity HardwareabstractWearable devices with built-in cameras present interesting opportunities for users to capture various aspects of their daily life and are potentially also useful in supporting users with low vision in their everyday tasks. However, state-of-the-art image wearables available in the market are limited to capturing images periodically and do not provide any real-time analysis of the data that might be useful for the wearers. In this paper, we present DeepEye - a match-box sized wearable camera that is capable of running multiple cloud-scale deep learn- ing models locally on the device, thereby enabling rich analysis of the captured images in near real-time without offloading them to the cloud. DeepEye is powered by a commodity wearable processor (Snapdragon 410) which ensures its wearable form factor. The software architecture for DeepEye addresses a key limitation with executing multiple deep learning models on constrained hardware, that is their limited runtime memory. We propose a novel inference software pipeline that targets the local execution of multiple deep vision models (specifically, CNNs) by interleaving the execution of computation-heavy convolutional layers with the loading of memory-heavy fully-connected layers. Beyond this core idea, the execution framework incorporates: a memory caching scheme and a selective use of model compression techniques that further minimizes memory bottlenecks. Through a series of experiments, we show that our execution framework outperforms the baseline approaches significantly in terms of inference latency, memory requirements and energy consumption. Akhil Mathur, Nicholas D. Lane, Sourav Bhattacharya, Aidan Boran, Claudio Forlivesi, Fahim Kawsar |
MobiSys | 2 |
| 2016 | Engagement-aware computing: modelling user engagement from mobile contextsabstractIn this paper, we examine the potential of using mobile context to model user engagement. Taking an experimental approach, we systematically explore the dynamics of user engagement with a smartphone through three different studies. Specifically, to understand the feasibility of detecting user engagement from mobile context, we first assess an EEG artifact with 10 users and observe a strong correlation between automatically detected engagement scores and user's subjective perception of engagement. Grounded on this result, we model a set of application level features derived from smartphone usage of 10 users to detect engagement of a usage session using a Random Forest classifier. Finally, we apply this model to train a variety of contextual factors acquired from smartphone usage logs of 130 users to predict user engagement using an SVM classifier with a F1-Score of 0.82. Our experimental results highlight the potential of mobile contexts in designing engagement-aware applications and provide guidance to future explorations. Akhil Mathur, Nicholas D. Lane, Fahim Kawsar |
UbiComp | 2 |
| 2016 | ppNav: Peer-to-Peer Indoor Navigation for SmartphonesabstractMost of existing indoor navigation systems work in a client/server manner, which needs to deploy comprehensive localization services together with precise indoor maps a prior. In this paper, we design and realize a Peer-to-Peer navigation system, named ppNav, on smartphones, which enables the fast-to-deploy navigation services, avoiding the requirements of pre-deployed location services and detailed floorplans. ppNav navigates a user to the destination by tracking user mobility, promoting timely walking tips, and alerting potential deviations, according to a previous traveller's trace experience. Specifically, we utilize the ubiquitous WiFi fingerprints in a novel diagrammed form and extract both radio and visual features of the diagram to track relative locations and exploit fingerprint similarity trend for deviation detection. Consolidating these techniques, we implement ppNav on commercial mobile devices and validate its performance in real environments. Our results show that ppNav achieves delightful performance, with an average relative error of 0.9m in trace tracking and a maximum delay of 9 samples (about 4.5s) in deviation detection. Zuwei Yin, Chenshu Wu, Zheng Yang 0002, Nicholas D. Lane, Yunhao Liu 0001 |
ICPADS | 4 |
| 2016 | HeadScan: A Wearable System for Radio-Based Sensing of Head and Mouth-Related ActivitiesabstractThe popularity of wearables continues to rise. However, their functionalities and applications are constrained by the types of sensors that are currently available. Accelerometers and gyroscopes struggle to capture complex user activities. Microphones and image sensors are more powerful but capture privacy sensitive information. Physiological sensors are obtrusive to users since they often require skin contact and must be placed at certain body positions to function. In contrast, radio- based sensing uses wireless radio signals to capture movements of different parts of body caused by human activities and therefore provides a contactless and privacy-preserving approach to detect and monitor human activities. In this paper, we contribute to the search for a new sensing modality for the next generation of wearable devices by exploring the feasibility of radio-based human activity sensing and recognition in the context of wearable setting. We envision radio-based sensing has the potential to fundamentally transform wearables as we currently know them. As the first step to achieve our vision, we have designed and developed HeadScan, a first- of-its-kind wearable for radio-based sensing of a number of human activities that involve head and mouth movements. HeadScan only requires a pair of small antennas placed on the shoulder and collar and one wearable unit worn on the arm or the belt of the user. HeadScan uses the fine-grained CSI measurements extracted from the radio signals and incorporates a radio signal processing pipeline that converts the raw CSI measurements into the targeted human activities. To examine the feasibility and performance of HeadScan, we have collected about 50.5 hours data from seven users. Our wide-range experiments including comparisons to a conventional skin-contact audio-based sensing approach to tracking the same set of head and mouth-related activities highlight the enormous potential of our radio-based sensing approach and provide guidance to future explorations. Biyi Fang, Nicholas D. Lane, Mi Zhang 0002, Fahim Kawsar |
IPSN | 2 |
| 2016 | DeepX: A Software Accelerator for Low-Power Deep Learning Inference on Mobile DevicesabstractBreakthroughs from the field of deep learning are radically changing how sensor data are interpreted to extract the high-level information needed by mobile apps. It is critical that the gains in inference accuracy that deep models afford become embedded in future generations of mobile apps. In this work, we present the design and implementation of DeepX, a software accelerator for deep learning execution. DeepX signif- icantly lowers the device resources (viz. memory, computation, energy) required by deep learning that currently act as a severe bottleneck to mobile adoption. The foundation of DeepX is a pair of resource control algorithms, designed for the inference stage of deep learning, that: (1) decompose monolithic deep model network architectures into unit- blocks of various types, that are then more efficiently executed by heterogeneous local device processors (e.g., GPUs, CPUs); and (2), perform principled resource scaling that adjusts the architecture of deep models to shape the overhead each unit-blocks introduces. Experiments show, DeepX can allow even large-scale deep learning models to execute efficently on modern mobile processors and significantly outperform existing solutions, such as cloud-based offloading. Nicholas D. Lane, Sourav Bhattacharya, Petko Georgiev, Claudio Forlivesi, Lei Jiao 0002, Lorena Qendro, Fahim Kawsar |
IPSN | 1 |
| 2016 | Demonstration Abstract: Accelerating Embedded Deep Learning Using DeepXabstractDeep learning has revolutionized the way sensor measurements are interpreted and application of deep learning has seen a great leap in inference accuracies in a number of fields. However, the significant requirement for memory and computational power has hindered the wide scale adoption of these novel computational techniques on resource constrained wearable and mobile platforms. In this demonstration we present DeepX, a software accelerator for efficiently running deep neural networks and convolutional neural networks on resource constrained embedded platforms, e.g., Nvidia Tegra K1 and Qualcomm Snapdragon 400. Nicholas D. Lane, Sourav Bhattacharya, Petko Georgiev, Claudio Forlivesi, Fahim Kawsar |
IPSN | 1 |
| 2016 | LEO: scheduling sensor inference algorithms across heterogeneous mobile processors and network resourcesabstractMobile apps that use sensors to monitor user behavior often employ resource heavy inference algorithms that make computational offloading a common practice. However, existing schedulers/offloaders typically emphasize one primary offloading aspect without fully exploring complementary goals (e.g., heterogeneous resource management with only partial visibility into underlying algorithms, or concurrent sensor app execution on a single resource) and as a result, may overlook performance benefits pertinent to sensor processing. Petko Georgiev, Nicholas D. Lane, Kiran Rachuri, Cecilia Mascolo |
MobiCom | 2 |
| 2016 | BodyScan: Enabling Radio-based Sensing on Wearable Devices for Contactless Activity and Vital Sign MonitoringabstractWearable devices are increasingly becoming mainstream consumer products carried by millions of consumers. However, the potential impact of these devices is currently constrained by fundamental limitations of their built-in sensors. In this paper, we introduce radio as a new powerful sensing modality for wearable devices and propose to transform radio into a mobile sensor of human activities and vital signs. We present BodyScan, a wearable system that enables radio to act as a single modality capable of providing whole-body continuous sensing of the user. BodyScan overcomes key limitations of existing wearable devices by providing a contactless and privacy-preserving approach to capturing a rich variety of human activities and vital sign information. Our prototype design of BodyScan is comprised of two components: one worn on the hip and the other worn on the wrist, and is inspired by the increasingly prevalent scenario where a user carries a smartphone while also wearing a wristband/smartwatch. This prototype can support daily usage with one single charge per day. Experimental results show that in controlled settings, BodyScan can recognize a diverse set of human activities while also estimating the user's breathing rate with high accuracy. Even in very challenging real-world settings, BodyScan can still infer activities with an average accuracy above 60% and monitor breathing rate information a reasonable amount of time during each day. Biyi Fang, Nicholas D. Lane, Mi Zhang 0002, Aidan Boran, Fahim Kawsar |
MobiSys | 2 |
| 2016 | Sparsification and Separation of Deep Learning Layers for Constrained Resource Inference on WearablesabstractDeep learning has revolutionized the way sensor data are analyzed and interpreted. The accuracy gains these approaches offer make them attractive for the next generation of mobile, wearable and embedded sensory applications. However, state-of-the-art deep learning algorithms typically require a significant amount of device and processor resources, even just for the inference stages that are used to discriminate high-level classes from low-level data. The limited availability of memory, computation, and energy on mobile and embedded platforms thus pose a significant challenge to the adoption of these powerful learning techniques. In this paper, we propose SparseSep, a new approach that leverages the sparsification of fully connected layers and separation of convolutional kernels to reduce the resource requirements of popular deep learning algorithms. As a result, SparseSep allows large-scale DNNs and CNNs to run efficiently on mobile and embedded hardware with only minimal impact on inference accuracy. We experiment using SparseSep across a variety of common processors such as the Qualcomm Snapdragon 400, ARM Cortex M0 and M3, and Nvidia Tegra K1, and show that it allows inference for various deep models to execute more efficiently; for example, on average requiring 11.3 times less memory and running 13.3 times faster on these representative platforms. Sourav Bhattacharya, Nicholas D. Lane |
SenSys | 2 |
| 2015 | Prime: a framework for co-located multi-device appsabstractEven though mobile devices are ubiquitous, the conceptually simple endeavor of using co-located devices for multi-user experiences is cumbersome. It may not even be possible when certain apps are not widely available. David Chu, Zengbin Zhang, Alec Wolman, Nicholas D. Lane |
UbiComp | 4 |
| 2015 | DeepEar: robust smartphone audio sensing in unconstrained acoustic environments using deep learningabstractMicrophones are remarkably powerful sensors of human behavior and context. However, audio sensing is highly susceptible to wild fluctuations in accuracy when used in diverse acoustic environments (such as, bedrooms, vehicles, or cafes), that users encounter on a daily basis. Towards addressing this challenge, we turn to the field of deep learning; an area of machine learning that has radically changed related audio modeling domains like speech recognition. In this paper, we present DeepEar -- the first mobile audio sensing framework built from coupled Deep Neural Networks (DNNs) that simultaneously perform common audio sensing tasks. We train DeepEar with a large-scale dataset including unlabeled data from 168 place visits. The resulting learned model, involving 2.3M parameters, enables DeepEar to significantly increase inference robustness to background noise beyond conventional approaches present in mobile devices. Finally, we show DeepEar is feasible for smartphones by building a cloud-free DSP-based prototype that runs continuously, using only 6% of the smartphone's battery daily. Nicholas D. Lane, Petko Georgiev, Lorena Qendro |
UbiComp | 1 |
| 2015 | More with less: lowering user burden in mobile crowdsourcing through compressive sensingabstractMobile crowdsourcing is a powerful tool for collecting data of various types. The primary bottleneck in such systems is the high burden placed on the user who must manually collect sensor data or respond in-situ to simple queries (e.g., experience sampling studies). In this work, we present Compressive CrowdSensing (CCS) -- a framework that enables compressive sensing techniques to be applied to mobile crowdsourcing scenarios. CCS enables each user to provide significantly reduced amounts of manually collected data, while still maintaining acceptable levels of overall accuracy for the target crowd-based system. Naïve applications of compressive sensing do not work well for common types of crowdsourcing data (e.g., user survey responses) because the necessary correlations that are exploited by a sparsifying base are hidden and non-trivial to identify. CCS comprises a series of novel techniques that enable such challenges to be overcome. We evaluate CCS with four representative large-scale datasets and find that it is able to outperform standard uses of compressive sensing, as well as conventional approaches to lowering the quantity of user data needed by crowd systems. Xiaohong Hao, Nicholas D. Lane, Xin Liu 0002, Thomas Moscibroda |
UbiComp | 3 |
| 2015 | Unobtrusive Sensing Incremental Social Contexts Using Fuzzy Class Incremental LearningabstractBy utilizing captured characteristics of surrounding contexts through widely used Bluetooth sensor, user-centric social contexts can be effectively sensed and discovered by dynamic Bluetooth information. At present, state-of-the-art approaches for building classifiers can basically recognize limited classes trained in the learning phase; however, due to the complex diversity of social contextual behavior, the built classifier seldom deals with newly appeared contexts, which results in degrading the recognition performance greatly. To address this problem, we propose, an OSELM (online sequential extreme learning machine) based class incremental learning method for continuous and unobtrusive sensing new classes of social contexts from dynamic Bluetooth data alone. We integrate fuzzy clustering technique and OSELM to discover and recognize social contextual behaviors by real-world Bluetooth sensor data. Experimental results show that our method can automatically cope with incremental classes of social contexts that appear unpredictably in the real-world. Further, our proposed method have the effective recognition capability for both original known classes and newly appeared unknown classes, respectively. Zhenyu Chen 0003, Yiqiang Chen 0001, Xingyu Gao 0001, Shuangquan Wang, Lisha Hu, Chenggang Yan 0001, Nicholas D. Lane, Chunyan Miao |
ICDM | 7 |
| 2015 | SIFT: building an internet of safe thingsabstractAs the number of connected devices explodes, the use scenarios of these devices and data have multiplied. Many of these scenarios, e.g., home automation, require tools beyond data visualizations, to express user intents and to ensure interactions do not cause undesired effects in the physical world. We present SIFT, a safety-centric programming platform for connected devices in IoT environments. First, to simplify programming, users express high-level intents in declarative IoT apps. The system then decides which sensor data and operations should be combined to satisfy the user requirements. Second, to ensure safety and compliance, the system verifies whether conflicts or policy violations can occur within or between apps. Through an office deployment, user studies, and trace analysis using a large-scale dataset from a commercial IoT app authoring platform, we demonstrate the power of SIFT and highlight how it leads to more robust and reliable IoT apps. Chieh-Jan Mike Liang, Börje Karlsson 0001, Nicholas D. Lane, Feng Zhao 0001, Junbei Zhang, Zheyi Pan, Yong Yu 0001 |
IPSN | 3 |
| 2015 | Cost-aware compressive sensing for networked sensing systemsabstractCompressive Sensing is a technique that can help reduce the sampling rate of sensing tasks. In mobile crowdsensing applications or wireless sensor networks, the resource burden of collecting samples is often a major concern. Therefore, compressive sensing is a promising approach in such scenarios. An implicit assumption underlying compressive sensing -- both in theory and its applications -- is that every sample has the same cost: its goal is to simply reduce the number of samples while achieving a good recovery accuracy. In many networked sensing systems, however, the cost of obtaining a specific sample may depend highly on the location, time, condition of the device, and many other factors of the sample. Xiaohong Hao, Nicholas D. Lane, Xin Liu 0002, Thomas Moscibroda |
IPSN | 3 |
| 2015 | ZOE: A Cloud-less Dialog-enabled Continuous Sensing Wearable Exploiting Heterogeneous ComputationabstractThe wearable revolution, as a mass-market phenomenon, has finally arrived. As a result, the question of how wearables should evolve over the next 5 to 10 years is assuming an increasing level of societal and commercial importance. A range of open design and system questions are emerging, for instance: How can wearables shift from being largely health and fitness focused to tracking a wider range of life events? What will become the dominant methods through which users interact with wearables and consume the data collected? Are wearables destined to be cloud and/or smartphone dependent for their operation? Nicholas D. Lane, Petko Georgiev, Cecilia Mascolo |
MobiSys | 1 |
| 2014 | Connecting personal-scale sensing and networked community behavior to infer human activitiesabstractAdvances in mobile and wearable devices are making it feasible to deploy sensing systems at a large-scale. However, slower progress is being made in activity recognition which remains often unreliable in everyday environments. In this paper, we investigate how to leverage the increasing capacity to gather data at a population-scale towards improving existing models of human behavior. Specifically, we consider the various social phenomena and environmental factors that cause people to develop correlated behavioral patterns, especially within communities connected by strong social ties. Reasons underpinning correlated behavior include shared externalities (e.g., work schedules, weather, traffic conditions), that shape options and decisions; and cases of adopted behavior, as people learn from each other or assume group norms due to social pressure. Most existing approaches to modeling human behavior ignore all of these phenomena and recognize activities solely on the basis of sensor data captured from a single individual. We propose the Networked Community Behavior (NCB) framework for activity recognition, specifically designed to exploit community-scale behavioral patterns. Under NCB, patterns of community behavior are mined to identify social ties that can signal correlated behavior, this information is used to augment sensor-based inferences available from the actions of individuals. Our evaluation of NCB shows it is able to outperform existing approaches to behavior modeling across four mobile sensing datasets that collectively require a diverse set of activities to be recognized. Nicholas D. Lane, Feng Zhao 0001 |
UbiComp | 1 |
| 2014 | Caiipa: automated large-scale mobile app testing through contextual fuzzingabstractScalable and comprehensive testing of mobile apps is extremely challenging. Every test input needs to be run with a variety of contexts, such as: device heterogeneity, wireless network speeds, locations, and unpredictable sensor inputs. The range of values for each context, e.g. location, can be very large. In this paper we present Caiipa, a cloud service for testing apps over an expanded mobile context space in a scalable way. It incorporates key techniques to make app testing more tractable, including a context test space prioritizer to quickly discover failure scenarios for each app. We have implemented Caiipa on a cluster of VMs and real devices that can each emulate various combinations of contexts for tablet and phone apps. We evaluate Caiipa by testing 265 commercially available mobile apps based on a comprehensive library of real-world conditions. Our results show that Caiipa leads to improvements of 11.1x and 8.4x in the number of crashes and performance bugs discovered compared to conventional UI-based automation (i.e., monkey-testing). Chieh-Jan Mike Liang, Nicholas D. Lane, Niels Brouwers, Börje Karlsson 0001, Hao Liu 0006, Xiang Shan, Ranveer Chandra, Feng Zhao 0001 |
MobiCom | 2 |
| 2014 | DSP.Ear: leveraging co-processor support for continuous audio sensing on smartphonesabstractThe rapidly growing adoption of sensor-enabled smartphones has greatly fueled the proliferation of applications that use phone sensors to monitor user behavior. A central sensor among these is the microphone which enables, for instance, the detection of valence in speech, or the identification of speakers. Deploying multiple of these applications on a mobile device to continuously monitor the audio environment allows for the acquisition of a diverse range of sound-related contextual inferences. However, the cumulative processing burden critically impacts the phone battery. Petko Georgiev, Nicholas D. Lane, Kiran Rachuri, Cecilia Mascolo |
SenSys | 2 |
| 2014 | BeWell: Sensing Sleep, Physical Activities and Social Interactions to Promote Wellbeing
Nicholas D. Lane, Mu Lin, Mashfiqui Mohammod, Xiaochao Yang, Hong Lu 0006, Giuseppe Cardone, Afsaneh Doryab, Ethan Berke, Andrew T. Campbell, Tanzeem Choudhury |
Mob. Networks Appl. | 1 |
| 2014 | Community Similarity Networks
Nicholas D. Lane, Hong Lu 0006, Shaohan Hu, Tanzeem Choudhury, Andrew T. Campbell, Feng Zhao 0001 |
Pers. Ubiquitous Comput. | 1 |
| 2013 | Understanding the coverage and scalability of place-centric crowdsensingabstractCrowd-enabled place-centric systems gather and reason over large mobile sensor datasets and target everyday user locations (such as stores, workplaces, and restaurants). Such systems are transforming various consumer services (for example, local search) and data-driven organizations (city planning). As the demand for these systems increases, our understanding of how to design and deploy successful crowdsensing systems must improve. In this paper, we present a systematic study of the coverage and scaling properties of place-centric crowdsensing. During a two-month deployment, we collected smartphone sensor data from 85 participants using a representative crowdsensing system that captures 48,000 different place visits. Our analysis of this dataset examines issues of core interest to place-centric crowdsensing, including place-temporal coverage, the relationship between the user population and coverage, privacy concerns, and the characterization of the collected data. Collectively, our findings provide valuable insights to guide the building of future place-centric crowdsensing systems and applications. Yohan Chon, Nicholas D. Lane, Yunjong Kim, Feng Zhao 0001, Hojung Cha |
UbiComp | 2 |
| 2013 | MoodScope: building a mood sensor from smartphone usage patternsabstractWe report a first-of-its-kind smartphone software system, MoodScope, which infers the mood of its user based on how the smartphone is used. Compared to smartphone sensors that measure acceleration, light, and other physical properties, MoodScope is a "sensor" that measures the mental state of the user and provides mood as an important input to context-aware computing. We run a formative statistical mood study with smartphone-logged data collected from 32 participants over two months. Through the study, we find that by analyzing communication history and application usage patterns, we can statistically infer a user's daily mood average with an initial accuracy of 66%, which gradu-ally improves to an accuracy of 93% after a two-month personal-ized training period. Motivated by these results, we build a service, MoodScope, which analyzes usage history to act as a sensor of the user's mood. We provide a MoodScope API for developers to use our system to create mood-enabled applications. We further create and deploy a mood-sharing social application. Robert LiKamWa, Yunxin Liu 0001, Nicholas D. Lane, Lin Zhong 0001 |
MobiSys | 3 |
| 2013 | MoodScope: building a mood sensor from smartphone usage patternsabstractWe present MoodScope, a software system which infers the mood of its user based on how the smartphone is used. Similar to smartphone sensors that measure acceleration, light, and other physical properties, MoodScope is a "sensor" that measures the mental state of the user and provides mood as an important input to context-aware computing. We run a formative statistical study with smartphone-logged data collected from 32 participants over two months. Through the study, we find that by analyzing communication history and application usage patterns, we can statistically infer a user's daily mood average with an accuracy of 93% after a two-month training period. Motivated by these results, we build a service, MoodScope, which analyzes usage history to act as a sensor of the user's mood. Robert LiKamWa, Yunxin Liu 0001, Nicholas D. Lane, Lin Zhong 0001 |
MobiSys | 3 |
| 2013 | CarSafe app: alerting drowsy and distracted drivers using dual cameras on smartphonesabstractWe present CarSafe, a new driver safety app for Android phones that detects and alerts drivers to dangerous driving conditions and behavior. It uses computer vision and machine learning algorithms on the phone to monitor and detect whether the driver is tired or distracted using the front-facing camera while at the same time tracking road conditions using the rear-facing camera. Today's smartphones do not, however, have the capability to process video streams from both the front and rear cameras simultaneously. In response, CarSafe uses acontext-aware algorithm that switches between the two cameras while processing the data in real-time with the goal of minimizing missed events inside (e.g., drowsy driving) and outside of the car (e.g., tailgating). Camera switching means that CarSafe technically has a "blind spot" in the front or rear at any given time. To address this, CarSafe uses other embedded sensors on the phone (i.e., inertial sensors) to generate soft hints regarding potential blind spot dangers. We present the design and implementation of CarSafe and discuss its evaluation using results from a 12-driver field trial. Results from the CarSafe deployment are promising -- CarSafe can infer a common set of dangerous driving behaviors and road conditions with an overall precision and recall of 83% and 75%, respectively. CarSafe is the first dual-camera sensing app for smartphones and represents a new disruptive technology because it provides similar advanced safety features otherwise only found in expensive top-end cars. Chuang-Wen You, Nicholas D. Lane, Rui Wang 0016, Zhenyu Chen 0003, Thomas J. Bao, Martha Montes-de-Oca, Yuting Cheng 0001, Mu Lin, Lorenzo Torresani, Andrew T. Campbell |
MobiSys | 2 |
| 2013 | CarSafe app: alerting drowsy and distracted drivers using dual cameras on smartphonesabstractWe present CarSafe, the first driver safety application that uses dual cameras on smartphones to detect and alert drivers to dangerous driving conditions. CarSafe fuses events detected from cameras and readings from embedded sensors on the phone -- such as the GPS, accelerometer and gyroscope -- to detect and alert the driver of dangerous driving behavior in and outside of the car. Results from a 12-driver field trial show CarSafe can infer five of the most commonly occurring dangerous driving conditions with an overall precision and recall of 83% and 75%, respectively. Chuang-Wen You, Nicholas D. Lane, Rui Wang 0016, Zhenyu Chen 0003, Thomas J. Bao, Martha Montes-de-Oca, Yuting Cheng 0001, Mu Lin, Lorenzo Torresani, Andrew T. Campbell |
MobiSys | 2 |
| 2013 | Piggyback CrowdSensing (PCS): energy efficient crowdsourcing of mobile sensor data by exploiting smartphone app opportunitiesabstractFueled by the widespread adoption of sensor-enabled smartphones, mobile crowdsourcing is an area of rapid innovation. Many crowd-powered sensor systems are now part of our daily life -- for example, providing highway congestion information. However, participation in these systems can easily expose users to a significant drain on already limited mobile battery resources. For instance, the energy burden of sampling certain sensors (such as WiFi or GPS) can quickly accumulate to levels users are unwilling to bear. Crowd system designers must minimize the negative energy side-effects of participation if they are to acquire and maintain large-scale user populations. Nicholas D. Lane, Yohan Chon, Yongzhe Zhang, Fan Li 0007, Guanzhong Ding, Feng Zhao 0001, Hojung Cha |
SenSys | 1 |
| 2012 | Towards Population Scale Activity Recognition: A Framework for Handling Data DiversityabstractThe rising popularity of the sensor-equipped smartphone is changing the possible scale and scope of human activity inference. The diversity in user population seen in large user bases can overwhelm conventional one-size-fits-all classification approaches. Although personalized models are better able to handle population diversity, they often require increased effort from the end user during training and are computationally expensive. In this paper, we propose an activity classification framework that is scalable and can tractably handle an increasing number of users. Scalability is achieved by maintaining distinct groups of similar users during the training process, which makes it possible to account for the differences between users without resorting to training individualized classifiers. The proposed framework keeps user burden low by leveraging crowd-sourced data labels, where simple natural language processing techniques in combination with multi-instance learning are used to handle labeling errors introduced by low-commitment everyday users. Experiment results on a large public dataset demonstrate that the framework can cope with population diversity irrespective of population size. Saeed Abdullah, Nicholas D. Lane, Tanzeem Choudhury |
AAAI | 2 |
| 2012 | Automatically characterizing places with opportunistic crowdsensing using smartphonesabstractAutomated and scalable approaches for understanding the semantics of places are critical to improving both existing and emerging mobile services. In this paper, we present [email protected] (CSP), a framework that exploits a previously untapped resource -- opportunistically captured images and audio clips from smartphones -- to link place visits with place categories (e.g., store, restaurant). CSP combines signals based on location and user trajectories (using WiFi/GPS) along with various visual and audio place "hints" mined from opportunistic sensor data. Place hints include words spoken by people, text written on signs or objects recognized in the environment. We evaluate CSP with a seven-week, 36-user experiment involving 1,241 places in five locations around the world. Our results show that CSP can classify places into a variety of categories with an overall accuracy of 69%, outperforming currently available alternative solutions. Yohan Chon, Nicholas D. Lane, Fan Li 0007, Hojung Cha, Feng Zhao 0001 |
UbiComp | 2 |
| 2012 | CarSafe demo: supporting driver safety using dual-cameras on smartphonesabstractWe demonstrate CarSafe, a driver safety application for Android phones that fuses information from both front and back cameras and others embedded sensors on the phone to detect and alert drivers to dangerous driving conditions in and outside of the car. In this demonstration, we set up an emulated driving environment to show how CarSafe works. Chuang-Wen You, Martha Montes-de-Oca, Thomas J. Bao, Nicholas D. Lane, Hong Lu 0006, Giuseppe Cardone, Lorenzo Torresani, Andrew T. Campbell |
UbiComp | 4 |
| 2012 | CarSafe: a driver safety app that detects dangerous driving behavior using dual-cameras on smartphonesabstractDriving while being tired or distracted is dangerous. We are developing the CafeSafe app for Android phones, which fuses information from both front and back cameras and others embedded sensors on the phone to detect and alert drivers to dangerous driving conditions in and outside of the car. CarSafe uses computer vision and machine learning algorithms on the phone to monitor and detect whether the driver is tired or distracted using the front camera while at the same time tracking road conditions using the back camera. CarSafe is the first dual-camera application for smart-phones. Chuang-Wen You, Martha Montes-de-Oca, Thomas J. Bao, Nicholas D. Lane, Hong Lu 0006, Giuseppe Cardone, Lorenzo Torresani, Andrew T. Campbell |
UbiComp | 4 |
| 2011 | Mobile sensing: challenges, opportunities and future directionsabstractThe emerging field of mobile sensing has engaged computer scientists from a variety of existing communities, such as, mobile systems, machine learning and human computer interaction. Each community approaches the challenges of mobile sensing research with its own unique perspective. The purpose of this workshop is to provide a forum to discuss the state of the art in mobile sensing and promote increased cooperation and interaction among the participating research communities. Nicholas D. Lane, Tanzeem Choudhury, Feng Zhao 0001 |
UbiComp | 1 |
| 2011 | Enabling large-scale human activity inference on smartphones using community similarity networks (csn)abstractSensor-enabled smartphones are opening a new frontier in the development of mobile sensing applications. The recognition of human activities and context from sensor-data using classification models underpins these emerging applications. However, conventional approaches to training classifiers struggle to cope with the diverse user populations routinely found in large-scale popular mobile applications. Differences between users (e.g., age, sex, behavioral patterns, lifestyle) confuse classifiers, which assume everyone is the same. To address this, we propose Community Similarity Networks (CSN), which incorporates inter-person similarity measurements into the classifier training process. Under CSN every user has a unique classifier that is tuned to their own characteristics. CSN exploits crowd-sourced sensor-data to personalize classifiers with data contributed from other similar users. This process is guided by similarity networks that measure different dimensions of inter-person similarity. Our experiments show CSN outperforms existing approaches to classifier training under the presence of population diversity. Nicholas D. Lane, Hong Lu 0006, Shaohan Hu, Tanzeem Choudhury, Andrew T. Campbell, Feng Zhao 0001 |
UbiComp | 1 |
| 2011 | Balancing energy, latency and accuracy for mobile sensor data classificationabstractSensor convergence on the mobile phone is spawning a broad base of new and interesting mobile applications. As applications grow in sophistication, raw sensor readings often require classification into more useful application-specific high-level data. For example, GPS readings can be classified as running, walking or biking. Unfortunately, traditional classifiers are not built for the challenges of mobile systems: energy, latency, and the dynamics of mobile. David Chu, Nicholas D. Lane, Tsung-Te Lai, Cong Pang, Xiangying Meng, Fan Li 0007, Feng Zhao 0001 |
SenSys | 2 |
| 2010 | Community-Guided Learning: Exploiting Mobile Sensor Users to Model Human BehaviorabstractModeling human behavior requires vast quantities of accurately labeled training data, but for ubiquitous people-aware applications such data is rarely attainable. Even researchers make mistakes when labeling data, and consistent, reliable labels from low-commitment users are rare. In particular, users may give identical labels to activities with characteristically different signatures (e.g., labeling eating at home or at a restaurant as "dinner") or may give different labels to the same context (e.g., "work" vs. "office"). In this scenario, labels are unreliable but nonetheless contain valuable information for classification. To facilitate learning in such unconstrained labeling scenarios, we propose Community-Guided Learning (CGL), a framework that allows existing classifiers to learn robustly from unreliably-labeled user-submitted data. CGL exploits the underlying structure in the data and the unconstrained labels to intelligently group crowd-sourced data. We demonstrate how to use similarity measures to determine when and how to split and merge contributions from different labeled categories and present experimental results that demonstrate the effectiveness of our framework. Daniel Peebles, Hong Lu 0006, Nicholas D. Lane, Tanzeem Choudhury, Andrew T. Campbell |
AAAI | 3 |
| 2010 | Hapori: context-based local search for mobile phones using community behavioral modeling and similarityabstractLocal search engines are very popular but limited. We present Hapori, a next-generation local search technology for mobile phones that not only takes into account location in the search query but richer context such as the time, weather and the activity of the user. Hapori also builds behavioral models of users and exploits the similarity between users to tailor search results to personal tastes rather than provide static geo-driven points of interest. We discuss the design, implementation and evaluation of the Hapori framework which combines data mining, information preserving embedding and distance metric learning to address the challenge of creating efficient multidimensional models from context-rich local search logs. Our experimental results using 80,000 queries extracted from search logs show that contextual and behavioral similarity information can improve the relevance of local search results by up to ten times when compared to the results currently provided by commercially available search engine technology. Nicholas D. Lane, Dimitrios Lymberopoulos, Feng Zhao 0001, Andrew T. Campbell |
UbiComp | 1 |
| 2010 | The Jigsaw continuous sensing engine for mobile phone applicationsabstractSupporting continuous sensing applications on mobile phones is challenging because of the resource demands of long-term sensing, inference and communication algorithms. We present the design, implementation and evaluation of the Jigsaw continuous sensing engine, which balances the performance needs of the application and the resource demands of continuous sensing on the phone. Jigsaw comprises a set of sensing pipelines for the accelerometer, microphone and GPS sensors, which are built in a plug and play manner to support: i) resilient accelerometer data processing, which allows inferences to be robust to different phone hardware, orientation and body positions; ii) smart admission control and on-demand processing for the microphone and accelerometer data, which adaptively throttles the depth and sophistication of sensing pipelines when the input data is low quality or uninformative; and iii) adaptive pipeline processing, which judiciously triggers power hungry pipeline stages (e.g., sampling the GPS) taking into account the mobility and behavioral patterns of the user to drive down energy costs. We implement and evaluate Jigsaw on the Nokia N95 and the Apple iPhone, two popular smartphone platforms, to demonstrate its capability to recognize user activities and perform long term GPS tracking in an energy-efficient manner. Hong Lu 0006, Zhigang Liu 0010, Nicholas D. Lane, Tanzeem Choudhury, Andrew T. Campbell |
SenSys | 4 |
| 2010 | Bubble-sensing: Binding sensing tasks to the physical world
Hong Lu 0006, Nicholas D. Lane, Shane B. Eisenman, Andrew T. Campbell |
Pervasive Mob. Comput. | 2 |
| 2009 | SoundSense: scalable sound sensing for people-centric applications on mobile phonesabstractTop end mobile phones include a number of specialized (e.g., accelerometer, compass, GPS) and general purpose sensors (e.g., microphone, camera) that enable new people-centric sensing applications. Perhaps the most ubiquitous and unexploited sensor on mobile phones is the microphone - a powerful sensor that is capable of making sophisticated inferences about human activity, location, and social events from sound. In this paper, we exploit this untapped sensor not in the context of human communications but as an enabler of new sensing applications. We propose SoundSense, a scalable framework for modeling sound events on mobile phones. SoundSense is implemented on the Apple iPhone and represents the first general purpose sound sensing system specifically designed to work on resource limited phones. The architecture and algorithms are designed for scalability and Soundsense uses a combination of supervised and unsupervised learning techniques to classify both general sound types (e.g., music, voice) and discover novel sound events specific to individual users. The system runs solely on the mobile phone with no back-end interactions. Through implementation and evaluation of two proof of concept people-centric sensing applications, we demostrate that SoundSense is capable of recognizing meaningful sound events that occur in users' everyday lives. Hong Lu 0006, Nicholas D. Lane, Tanzeem Choudhury, Andrew T. Campbell |
MobiSys | 3 |
| 2009 | BikeNet: A mobile sensing system for cyclist experience mappingabstractWe present BikeNet, a mobile sensing system for mapping the cyclist experience. Built leveraging the MetroSense architecture to provide insight into the real-world challenges of people-centric sensing, BikeNet uses a number of sensors embedded into a cyclist's bicycle to gather quantitative data about the cyclist's rides. BikeNet uses a dual-mode operation for data collection, using opportunistically encountered wireless access points in a delay-tolerant fashion by default, and leveraging the cellular data channel of the cyclist's mobile phone for real-time communication as required. BikeNet also provides a Web-based portal for each cyclist to access various representations of her data, and to allow for the sharing of cycling-related data (for example, favorite cycling routes) within cycling interest groups, and data of more general interest (for example, pollution data) with the broader community. We present: a description and prototype implementation of the system architecture based on customized Moteiv Tmote Invent motes and sensor-enabled Nokia N80 mobile phones; an evaluation of sensing and inference that quantifies cyclist performance and the cyclist environment; a report on networking performance in an environment characterized by bicycle mobility and human unpredictability; and a description of BikeNet system user interfaces. Shane B. Eisenman, Emiliano Miluzzo, Nicholas D. Lane, Ronald A. Peterson, Gahng-Seop Ahn, Andrew T. Campbell |
ACM Trans. Sens. Networks | 3 |
| 2008 | Techniques for Improving Opportunistic Sensor Networking Performance
Shane B. Eisenman, Nicholas D. Lane, Andrew T. Campbell |
DCOSS | 2 |
| 2008 | CaliBree: A Self-calibration System for Mobile Sensor Networks
Emiliano Miluzzo, Nicholas D. Lane, Andrew T. Campbell, Reza Olfati-Saber |
DCOSS | 2 |
| 2008 | Transforming the social networking experience with sensing presence from mobile phones
Andrew T. Campbell, Shane B. Eisenman, Kristóf Fodor, Nicholas D. Lane, Hong Lu 0006, Emiliano Miluzzo, Mirco Musolesi, Ronald A. Peterson |
SenSys | 4 |
| 2008 | Sensing meets mobile social networks: the design, implementation and evaluation of the cenceme applicationabstractWe present the design, implementation, evaluation, and user ex periences of theCenceMe application, which represents the first system that combines the inference of the presence of individuals using off-the-shelf, sensor-enabled mobile phones with sharing of this information through social networking applications such as Facebook and MySpace. We discuss the system challenges for the development of software on the Nokia N95 mobile phone. We present the design and tradeoffs of split-level classification, whereby personal sensing presence (e.g., walking, in conversation, at the gym) is derived from classifiers which execute in part on the phones and in part on the backend servers to achieve scalable inference. We report performance measurements that characterize the computational requirements of the software and the energy consumption of the CenceMe phone client. We validate the system through a user study where twenty two people, including undergraduates, graduates and faculty, used CenceMe continuously over a three week period in a campus town. From this user study we learn how the system performs in a production environment and what uses people find for a personal sensing system. Emiliano Miluzzo, Nicholas D. Lane, Kristóf Fodor, Ronald A. Peterson, Hong Lu 0006, Mirco Musolesi, Shane B. Eisenman, Andrew T. Campbell |
SenSys | 2 |
| 2008 | Integrating sensor presence into virtual worlds using mobile phonesabstractNo abstract available. Mirco Musolesi, Emiliano Miluzzo, Nicholas D. Lane, Shane B. Eisenman, Tanzeem Choudhury, Andrew T. Campbell |
SenSys | 3 |
| 2007 | The BikeNet mobile sensing system for cyclist experience mappingabstractWe describe our experiences deploying BikeNet, an extensible mobile sensing system for cyclist experience mapping leveraging opportunistic sensor networking principles and techniques. BikeNet represents a multifaceted sensing system and explores personal, bicycle, and environmental sensing using dynamically role-assigned bike area networking based on customized Moteiv Tmote Invent motes and sensor-enabled Nokia N80 mobile phones. We investigate real-time and delay-tolerant uploading of data via a number of sensor access points (SAPs) to a networked repository. Among bicycles that rendezvous en route we explore inter-bicycle networking via data muling. The repository provides a cyclist with data archival, retrieval, and visualization services. BikeNet promotes the social networking of the cycling community through the provision of a web portal that facilitates back end sharing of real-time and archived cycling-related data from the repository. We present: a description and prototype implementation of the system architecture, an evaluation of sensing and inference that quantifies cyclist performance and the cyclist environment; a report on networking performance in an environment characterized by bicycle mobility and human unpredictability; and a description of BikeNet system user interfaces. Visit [4] to see how the BikeNet system visualizes a user's rides. Shane B. Eisenman, Emiliano Miluzzo, Nicholas D. Lane, Ronald A. Peterson, Gahng-Seop Ahn, Andrew T. Campbell |
SenSys | 3 |
| 2006 | Virtual sensing range
Emiliano Miluzzo, Nicholas D. Lane, Andrew T. Campbell |
SenSys | 2 |