Sasu Tarkoma

dblp:66/5629 · DBLP profile ↗
← Back
160ranked-venue papers
8as first author
58since 2021 · last 2026
0000-0003-4220-3650ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 75 · 5 first-author · 28 since 2021Databases, data management, data science and information retrieval · 23 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 21 · 12 since 2021Human-computer interaction and ubiquitous computing · 17 · 1 first-author · 4 since 2021Systems, architecture and hardware · 16 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 11 · 4 since 2021Software engineering, systems software and programming languages · 5 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 since 2021Security and privacy · 4 · 2 since 2021
YearPublicationVenuePosition
2026 BridgeLoRA: Privacy-Preserving Collaborative Skip-Layer Connectors for Efficient Transformer Fine-Tuning at the Edge
abstract
fi=vertaisarvioitu|en=peerReviewed|
Vilhelm Toivonen, Xiang Su 0001, Xiaoli Liu 0005, Sasu Tarkoma, Pan Hui 0001
ICDCS4
2026 HandoffFL: Efficient Communication in Federated Learning Via Asynchronous Layer-Wise Handoff Training
abstract
Federated Learning (FL) enables collaborative training without sharing raw data, yet conventional synchronous approaches suffer from poor scalability and high communication costs, imposing heavy burdens on resource-constrained clients in heterogeneous edge environments. While asynchronous FL reduces heterogeneous stragglers, it introduces severe staleness and convergence degradation. Recent layer-wise FL methods reduce resource consumption by decomposing training into sequential layer-level tasks; however, they still rely on rigid round-based synchronization and therefore lack adaptability to dynamic client availability and network fluctuations. We propose HandoffFL, an event-driven asynchronous layer-wise FL framework built on a real-time message-passing monitoring architecture. HandoffFL organizes training into layer-level stages, where clients train one layer at a time with frozen prefixes. Transitions between layers are triggered by adaptive convergence events, enabling clients to progressively hand off converged layers and advance at their own pace. To support convergence detection at intermediate layers, the framework introduces temporal classifiers and further applies layer-specific staleness-aware aggregation with exponential decay and truncation. Extensive experiments on MNIST, Fashion-MNIST, and CIFAR-10 with multiple model architectures demonstrate that HandoffFL substantially reduces wall-clock training time and communication overhead compared to synchronous layer-wise FL, while maintaining comparable accuracy and robustness under heterogeneous client conditions compared with asynchronous FL.
Wenjun Zhang 0013, Sasu Tarkoma
ICDCS3
2026 Digital Twins for smart campus networks: An end-to-end framework for multi-domain data intelligence
abstract
Heterogeneous communication environments are expected to be pivotal in sixth-generation (6G) networks, enabling devices to seamlessly utilize various coexisting connectivity options, including cellular technologies (e.g., 4G/5G) and complementary access technologies like Wi-Fi, Bluetooth, and other short-range wireless systems. The challenge lies not just in device diversity but in effectively managing environments where nodes can dynamically switch between different network interfaces to optimize performance, reliability, and quality of service. This paper explores the role of Digital Twins (DTs) as an enabling framework for managing these multi-connectivity environments, where both mobile and non-cellular access technologies collectively enhance overall network functionality. We highlight key limitations in current DT implementations for these connectivity-rich scenarios and propose a scalable DT framework tailored explicitly for integrated cellular and auxiliary wireless access. Developed with commercial mobile devices as endpoints and a widely adopted twinning platform, our solution enhances management efficiency, optimizes resource allocation, and improves Quality of Service (QoS) in dynamic connectivity settings. We evaluate two architectural designs for DT deployment with respect to workload distribution and scalability. Additionally, we present a generative AI-driven analytics pipeline that consolidates descriptive, diagnostic, predictive, and prescriptive analytics into a unified closed-loop system. Experimental results from a real-world smart campus environment show that the offloaded architecture keeps CPU usage below 6% on mobile devices with up to 40 nodes while maintaining low latency and stable performance. The analytics pipeline achieves near-interactive response times, ensuring real-time network awareness and adaptive decision-making in multi-connectivity network environments.
Bivek Pandey, Yasith R. Wanigarathna, Sasu Tarkoma, Roberto Morabito
Comput. Commun.3
2026 SARF: Sparsity-Aware Reconstruction Framework for Large-Scale Datasets
abstract
Large-scale datasets, particularly those collected from smart devices and Internet of Things sensors, usually exhibit significant temporal and spatial sparsity, resulting in high amounts of missing data. Unless addressed in the analysis, this sparsity can result in substantial gaps and biases as well as limit the generalizability of conclusions drawn from such data. To address this challenge in data quality, we contribute the Sparsity-Aware Reconstruction Framework (SARF) as a novel and unified data fusion and reconstruction framework that enhances data quality and addresses sparsity. SARF analyzes datasets, partitioning the data into segments with similar characteristics, and reconstructs the data in each segment individually by selecting a reconstruction technique that is tailored to the internal temporal-spatial characteristics of the dataset. Through extensive experiments on two representative datasets - mobile application measurements and IoT sensor data from low-cost air quality sensors - we demonstrate that the targeted adaptation of reconstruction strategies employed by SARF significantly enhances the quality of reconstructed data. Our results show the robustness of SARF's performance across spatiotemporal variations, outperforming current state-of-the-art methods by margins up to 68% on average (74% for compressive sensing, 53% for convolutional sparse coding, 78% for deep learning). These findings underscore SARF's potential to enhance datadriven insights across multiple domains, paving the way for more robust analyses of sparsity-affected datasets.
Agustin Zuniga, Huber Flores, Ngoc Thi Nguyen, Pan Hui 0001, Sasu Tarkoma, Petteri Nurmi
IEEE Trans. Big Data5
2026 Sometimes Painful but Promising: Feasibility and Trade-Offs of On-Device Language Model Inference
abstract
The rapid rise of Language Models (LMs) has expanded the capabilities of natural language processing, powering applications from text generation to complex decision-making. While state-of-the-art LMs often boast hundreds of billions of parameters and are primarily deployed in data centers, recent trends show a growing focus on compact models—typically under 10 billion parameters–enabled by techniques such as quantization and other model compression techniques. This shift paves the way for LMs on edge devices, offering potential benefits such as enhanced privacy, reduced latency, and improved data sovereignty. However, the inherent complexity of even these smaller models, combined with the limited computing resources of edge hardware, raises critical questions about the practical trade-offs in executing LM inference outside the cloud. To address these challenges, we present a comprehensive evaluation of generative LM inference on representative CPU-based and GPU-accelerated edge devices. Our study measures key performance indicators—including memory usage, inference speed, and energy consumption—across various device configurations. Additionally, we examine throughput-energy trade-offs, cost considerations, and usability, alongside an assessment of qualitative model performance. While quantization helps mitigate memory overhead, it does not fully eliminate resource bottlenecks, especially for larger models. Our findings quantify the memory and energy constraints that must be considered for practical real-world deployments, offering concrete insights into the trade-offs between model size, inference performance, and efficiency. The exploration of LMs at the edge is still in its early stages. We hope this study provides a foundation for future research, guiding the refinement of models, the enhancement of inference efficiency, and the advancement of edge-centric AI systems.
Maximilian Abstreiter, Sasu Tarkoma, Roberto Morabito
ACM Trans. Embed. Comput. Syst.2
2025 Exponentially Weighted Instance-Aware Repeat Factor Sampling for Long-Tailed Object Detection Model Training in Unmanned Aerial Vehicles Surveillance Scenarios
abstract
Object detection models often struggle with class imbalance, where rare categories appear significantly less frequently than common ones. Existing sampling-based rebalancing strategies, such as Repeat Factor Sampling (RFS) and Instance-Aware Repeat Factor Sampling (IRFS), mitigate this issue by adjusting sample frequencies based on image and instance counts. However, these methods are based on linear adjustments, which limit their effectiveness in long-tailed distributions. This work introduces Exponentially Weighted Instance-Aware Repeat Factor Sampling (E-IRFS), an extension of IRFS that applies exponential scaling to better differentiate between rare and frequent classes. E-IRFS adjusts sampling probabilities using an exponential function applied to the geometric mean of image and instance frequencies, ensuring a more adaptive rebalancing strategy. We evaluate E-IRFS on a dataset derived from the Fireman-UAV-RGBT Dataset and four additional public datasets, using YOLOv11 object detection models to identify fire, smoke, people and lakes in emergency scenarios. The results show that E-IRFS improves detection performance by 22% over the baseline and outperforms RFS and IRFS, particularly for rare categories. The analysis also highlights that E-IRFS has a stronger effect on lightweight models with limited capacity, as these models rely more on data sampling strategies to address class imbalance. The findings demonstrate that E-IRFS improves rare object detection in resource-constrained environments, making it a suitable solution for real-time applications such as UAV-based emergency monitoring. The code is available at: https://github.com/futurians/E-IRFS.
Taufiq Ahmed, Abhishek Kumar 0011, Constantino Álvarez Casado, Anlan Zhang, Tuomo Hänninen, Lauri Lovén, Miguel Bordallo López, Sasu Tarkoma
IROS8
2025 Poster: IoT-Based Indoor Air Quality Monitoring for Health Risk Assessment and Well-Being
abstract
We study the impact of air pollutants on occupants' health by deploying 26 low-cost IoT air quality sensors in the UbiKampus office space at the University of Helsinki. Our data shows significant air quality variation even within meter-scale distances, highlighting the need for detailed indoor air quality monitoring and emphasizing the potential benefits low-cost IoT sensors can bring.
Naser Hossein Motlagh, Martha Arbayani Zaidan, Pak Lun Fung, Samu Varjonen, Andrew Rebeiro-Hargrave, Petteri Nurmi, Sasu Tarkoma
MobiSys7
2025 Drones in the Sky: Air Quality Monitoring at Heights
abstract
Air pollution represents a critical global health challenge. Traditional air quality monitoring methods, which rely on expensive stations or extensive low-cost sensor networks, often fall short in capturing the fine spatial and temporal variations of pollutants, particularly in urban environments dominated by vehicular emissions. Recent advancements in drone technology offer a novel solution to these limitations, enabling the collection of high-resolution air quality data across different altitudes and environments. We contribute by investigating the potential of drone-based air quality monitoring. Specifically, we conduct measurements in two distinct settings: an industrial site and a residential area. We present analytical findings that highlight the effectiveness of drones in capturing pollutant levels and discuss the key challenges associated with this technology. Our findings underscore the promise of drone-based monitoring in enhancing air quality assessment and inform directions for future research.
Naser Hossein Motlagh, Martha Arbayani Zaidan, Matti Irjala, Andrew Rebeiro-Hargrave, Petteri Nurmi, Sasu Tarkoma
MobiSys6
2025 Dynamic Hierarchical Reinforcement Learning Framework for Energy-Efficient 5G Base Stations in Urban Environments
abstract
The energy consumption of 5G base stations (BSs) is significantly higher than that of 4G BSs, creating challenges for operators due to increased costs and carbon emissions. Existing solutions address this issue by switching off BSs during specific periods or forming cooperation coalitions where some BSs deactivate while others serve users. However, these approaches often rely on fixed geographic configurations, making them unsuitable for urban areas with numerous BSs and mobile users. To tackle these challenges, we propose a hierarchical reinforcement learning (RL) framework for energy conservation in large-scale 5G networks. In the upper-layer, we propose a deep Q-network integrated with a graph convolutional network that dynamically groups BSs into coalitions from a macro perspective. This layer focuses on high-level coalition formation to optimize system-wide energy efficiency by considering the global state of the network. In the lower-layer, we combine attention mechanism with multi-agent RL and graph convolutional networks to design a scalable algorithm that maximizes local energy efficiency through optimizing the cooperation within each coalition. These two layers align global coalition dynamics with local intra-coalition cooperation to achieve system-wide energy optimization. Moreover, we accurately model large-scale urban 5G scenarios leveraging a high-fidelity network simulator, which enables our RL framework to learn from real-world feedback. Extensive experiments conducted with the simulator demonstrate that our proposed framework achieves remarkable energy savings of up to 75.6%, significantly outperforming baseline approaches. These findings highlight the effectiveness and superiority of our hierarchical RL optimization framework in addressing the energy consumption challenges faced by large-scale 5G networks.
Dianlei Xu, Xiang Su 0001, Gopika Premsankar, Huandong Wang, Sasu Tarkoma, Pan Hui 0001
IEEE Trans. Mob. Comput.5
2025 FPSelector: A Flexible Path Selector for Mobile Augmented Reality Offloading
abstract
Mobile Augmented Reality (MAR) applications pose unique challenges due to computation intensity, constrained device resources, and high interactive rendering requirements. The emergence of 5 G and edge computing offers opportunities to offload computation to the edge and cloud, indirectly enhancing the computing capability and usage duration of MAR devices. However, existing general task offloading and multipath transmission techniques do not address the challenges in offloading path selection with multiple edges, dynamic resource competition awareness, and spatial computation with strong task dependencies. This paper contributes FPSelector, a flexible path selector for MAR offloading. We present a two-tier MAR-specific offloading scheme with multiple edge nodes. In offloading decisions, we design a reinforcement learning model to generate the selection policy for each packet of an AR data stream. This model incorporates an action masking mechanism, a comprehensive reward function, and state features complemented by a resource prediction module, making FPSelector aware of dynamic heterogeneous environments. Moreover, we propose an online learning strategy to facilitate real-time selection. To validate its efficacy, we compare FPSelector's performance against leading schedulers under various scenarios, demonstrating a notable reduction of 9.9% and 9.6% in overall completion time for 4 K and 8 K video-based MAR applications compared to its closest competitor.
Yuanwei Zhu, Yakun Huang, Xiuquan Qiao, Xiaoli Liu 0005, Xiang Su 0001, Anna Brunström, Özgü Alay, Sasu Tarkoma
IEEE Trans. Mob. Comput.8
2025 HeLoRA: LoRA-heterogeneous Federated Fine-tuning for Foundation Models
abstract
Foundation models (FMs) have achieved state-of-the-art performance across various domains, benefiting from their vast number of parameters and the extensive amount of publicly available training data. However, real-world deployments reveal challenges such as system heterogeneity, where not all devices can handle the complexity of FMs, and emerging privacy concerns that limit the availability of public data. To address these challenges, we propose HeLoRA, a novel approach combining low-rank adaptation (LoRA) with federated learning to enable heterogeneous federated fine-tuning. HeLoRA allows clients to fine-tune models with different complexities by adjusting the rank values of LoRA matrices, tailoring the process to each device’s capabilities. To tackle the challenge of aggregating models with different structures, HeLoRA introduces two variants, i.e., HeLoRA-Pad and HeLoRA-KD. HeLoRA-Pad employs context-based padding to standardize the LoRA matrices, aligning them with the global model through a rank-based adaptive aggregation strategy. In contrast, HeLoRA-KD leverages the idea of deep mutual learning for aggregation, allowing heterogeneous models to retain their original structures. Extensive experiments with various datasets and ablation studies demonstrate that HeLoRA outperforms existing baselines, promising to enhance the practical deployment of FMs in diverse real-world environments.
Boyu Fan, Xiang Su 0001, Sasu Tarkoma, Pan Hui 0001
ACM Trans. Internet Techn.3
2024 A Survey on Model-heterogeneous Federated Learning: Problems, Methods, and Prospects
abstract
As privacy concerns continue to grow, federated learning (FL) has gained significant attention as a promising privacy-preserving technology, leading to considerable advancements in recent years. Unlike traditional machine learning, which requires central data collection, FL keeps data localized on user devices. However, conventional FL assumes that all clients operate with identical model structures initialized by the server. In real-world applications, system heterogeneity is common, with clients possessing varying computational capabilities. This disparity can hinder training for resource-limited clients and result in inefficient resource use for those with greater processing power. To address this challenge, model-heterogeneous FL has been introduced, enabling clients to train models of varying complexity based on their hardware resources. This paper reviews state-of-the-art approaches in model-heterogeneous FL, analyzing their strengths and weaknesses, while identifying open challenges and future research directions. To the best of our knowledge, this is the first survey to specifically focus on model-heterogeneous FL.
Boyu Fan, Siyang Jiang, Xiang Su 0001, Sasu Tarkoma, Pan Hui 0001
IEEE Big Data4
2024 PhD School: Network-Enhanced On-Device AI: a Recipe for Interoperable and Cooperative AI Applications
Paulius Daubaris, Sasu Tarkoma, Roberto Morabito
EWSN2
2024 GeoFed: A Geography-Aware Federated Learning Approach for Vehicular Visual Crowdsensing
abstract
Internet of Things (IoT) technology enables enhanced connectivity and information sharing among various devices and platforms. In the context of vehicular crowds ensing, this connectivity has opened up new way to collect environmental data via Vehicle-based Visual Crowdsensing. However, the heterogeneity of data sources and the presence of vehicle outliers pose challenges of ensuring the reliability and accuracy of the machine learning (ML) models. We propose GeoFed, a geography-aware federated learning (FL) approach for vehicular visual crowdsensing. Here, geographically similar vehicular fog nodes (VFNs) collaborate to train a cluster model unlike the traditional FL approaches where vehicles participate to train a model. To further improve GeoFed's performance, we employ the deep Q-Network (DQN) algorithm to intelligently determine the participation of vehicles in the FL process. Through extensive experiments on our own collected real-world dataset, we find that our proposed GeoFed not only outperforms the state-of-art FedAvg with higher F1 score (1.18 x) and mAP (1.14 x), but also achieves a faster convergence rate with less loss (80%).
Xinli Hao, Wenjun Zhang 0013, Xiaoli Liu 0001, Chao Zhu 0002, Sasu Tarkoma
ICC5
2024 From Pixels to Progress: Generating Road Network from Satellite Imagery for Socioeconomic Insights in Impoverished Areas
Yanxin Xi, Yu Liu 0016, Sasu Tarkoma, Pan Hui 0001, Yong Li 0008
IJCAI4
2024 Future of software development with generative AI
abstract
Abstract Generative AI is regarded as a major disruption to software development. Platforms, repositories, clouds, and the automation of tools and processes have been proven to improve productivity, cost, and quality. Generative AI, with its rapidly expanding capabilities, is a major step forward in this field. As a new key enabling technology, it can be used for many purposes, from creative dimensions to replacing repetitive and manual tasks. The number of opportunities increases with the capabilities of large-language models (LLMs). This has raised concerns about ethics, education, regulation, intellectual property, and even criminal activities. We analyzed the potential of generative AI and LLM technologies for future software development paths. We propose four primary scenarios, model trajectories for transitions between them, and reflect against relevant software development operations. The motivation for this research is clear: the software development industry needs new tools to understand the potential, limitations, and risks of generative AI, as well as guidelines for using it.
Jaakko J. Sauvola, Sasu Tarkoma, Mika Klemettinen, Jukka Riekki, David S. Doermann
Autom. Softw. Eng.2
2024 SimCost: cost-effective resource provision prediction and recommendation for spark workloads
abstract
Abstract Spark is one of the most popular big data analytical platforms. To save time, achieve high resource utilization, and remain cost-effective for Spark jobs, it is challenging but imperative for data scientists to configure suitable resource portions.In this paper, we investigate the proper parameter values that meet workloads’ performance requirements with minimized resource cost and resource utilization time. We propose SimCost , a simulation-based cost model, to predict the performance of jobs accurately. We achieve low-cost training by taking advantage of simulation framework , i.e., Monte Carlo simulation, which uses a small amount of data and resources to make a reliable prediction for larger datasets and clusters. Our method’s salient feature is that it allows us to invest low training costs while obtaining an accurate prediction. Through empirical experiments with 12 benchmark workloads, we show that the cost model yields less than 5% error on average prediction accuracy, and the recommendation achieves up to 6x resource cost saving.
Yuxing Chen 0003, Mohammad Ashraful Hoque, Pengfei Xu 0004, Jiaheng Lu, Sasu Tarkoma
Distributed Parallel Databases5
2024 Access control for trusted data sharing
abstract
Abstract In the envisioned 6G landscape, data sharing is expected to become increasingly prevalent, giving rise to digital marketplaces that foster cooperation among organizations for collecting, sharing, and processing data for analysis. These marketplaces serve as connectors between data producers and consumers, empowering multi-tenancy scenarios for seamless and secure data sharing both within and outside organizations. Given that 6G networks promise ultra-low latency, enhanced connectivity, and massive data throughput, the need for robust data access control mechanisms becomes imperative. These mechanisms ensure security and trust among entities, particularly in multi-tenant environments where multiple organizations share infrastructure and data resources. In this paper, we have designed and implemented a novel access control mechanism tailored for a distributed data streaming system developed by Nokia Bell Labs. Our approach leverages fine-grained policies, dynamic enforcement, and transparency mechanisms to enhance trust between data owners and consumers. By facilitating secure multi-tenancy data sharing, our solution contributes to the seamless exchange of data across diverse entities within the next-generation communication ecosystem. We demonstrate that our proposed access control mechanism incurs minimal overhead while ensuring data confidentiality and integrity. The introduction of such advancements in data sharing markets strengthens the overall ecosystem by providing heightened transparency and enhanced control over data, promoting collaboration and innovation in the 6G era.
Maria Zubair, Maryam Sabzevari, Vikramajeet Khatri, Sasu Tarkoma, Kimmo Hätönen
EURASIP J. Inf. Secur.4
2024 Estimating Black Carbon Levels With Proxy Variables and Low-Cost Sensors
abstract
We develop a portable and affordable solution for estimating personal exposure to black carbon (BC) using low-cost sensors and machine learning. Our approach uses other pollutants and environmental variables as proxies for estimating the concentrations of BC and combines this with machine learning based sensor calibration to improve the quality of the inputs that are used as proxies in the modeling. We extensively validate the feasibility of our approach and demonstrate its benefits with benchmarks conducted on real world data from two different urban locations with different population densities and characteristics. Our results demonstrate that our approach can accurately estimate BC (R2 higher than 0.9) without relying on a dedicated sensor. The results also highlight how calibration is essential for ensuring accurate modeling on low-cost sensor measurements. Our results offer a novel affordable and portable solution that can be used to estimate personal exposure to BC and, more generally, demonstrate how low-cost sensors and proxy modeling can increase the spatiotemporal scale at which information about BC level is available.
Xiaoli Liu 0005, Francesco Concas, Naser Hossein Motlagh, Martha Arbayani Zaidan, Pak Lun Fung, Samu Varjonen, Jarkko V. Niemi, Hilkka Timonen, Tareq Hussein, Tuukka Petäjä, Markku Kulmala, Petteri Nurmi, Sasu Tarkoma
IEEE Internet Things J.13
2024 Digital Twins for Smart Spaces - Beyond IoT Analytics
abstract
Smart spaces, physical spaces that are integrated with sensor-enabled IoT devices, are a powerful paradigm for optimizing the operations of the space and improving its quality for the occupants. Managing the applications and services running in the space is a complex task as the operations of the devices and services are dependent on the physical characteristics of the space, the occupants of the space, and the technologies that are being integrated. Digital twinning, the combination of physical representations with a virtual counterpart, is a potential technology for facilitating the management of smart space devices and services. While digital twins are increasingly adopted in industry, their use in everyday environments remains low due to difficulties in creating and linking the virtual representation with the physical environment. In this paper, we propose our vision for the adoption of digital twinning as a pathway to improve the functions of smart spaces. We derive a generic reference architecture that comprises four layers, covering the physical space, the sensing infrastructure, the network interfaces, and the underlying computational infrastructure. Next, we identify and address key requirements for the uptake of digital twins in smart space and assess their benefits using the ascendancy model of business analytics. Finally, to demonstrate the practicality of digital twinning, we present a proof-of-concept digital twin for the TellUs smart space at the University of Oulu in Finland and use it to highlight the potential benefits of different ascendancy levels.
Naser Hossein Motlagh, Martha Arbayani Zaidan, Lauri Lovén, Pak Lun Fung, Tuomo Hänninen, Roberto Morabito, Petteri Nurmi, Sasu Tarkoma
IEEE Internet Things J.8
2024 FedVisual: Heterogeneity-Aware Model Aggregation for Federated Learning in Visual-Based Vehicular Crowdsensing
abstract
With the advancement of assisted and autonomous driving technologies, vehicles are being outfitted with an ever-increasing number of sensors. Among these, visible light sensors, or dash-cameras, produce visual data rich in information. Analyzing this visual data through crowdsensing allows for low-cost and timely perception of urban road conditions, such as identifying dangerous driving behaviors and locating parking spaces. However, uploading such massive visual data to the cloud for centralized processing can lead to significant bandwidth challenges and also raise privacy concerns among vehicle owners. Federated learning (FL), in which vehicles serve as both data generators and computing nodes, presents a promising solution to address these challenges. Nevertheless, urban roads are complex and vehicles in different locations encounter completely different scenes, resulting in non-independently and identically distributed (non-i.i.d.) characteristics. Additionally, the diversity in dash-camera and onboard computation resources may lead to differences in the performance of locally trained models. Indiscriminate aggregating of local models from all vehicles can potentially degrade the global model’s performance. To overcome these challenges, we introduce FedVisual, a model aggregation approach for FL in vehicular visual crowdsensing. FedVisual leverages deep Q-network (DQN) to select appropriate local models, considering the heterogeneities in visual data contents and vehicles’ specifications. By leveraging the historical training experience, an effective model selection strategy can be obtained without complex mathematical modeling. Through the extensive simulations of our self-collected driving videos, FedVisual reduces model aggregation latency by up to 3.8% while improving the model’s performance by up to 3.2% compared to reference works.
Wenjun Zhang 0013, Xiaoli Liu 0005, Ruoyi Zhang, Chao Zhu 0002, Sasu Tarkoma
IEEE Internet Things J.5
2024 An Urban Trajectory Data-Driven Approach for COVID-19 Simulation
abstract
The coronavirus disease 2019 (COVID-19) pandemic has changed the world deeply. Urban trajectory big data collected by wireless sensing devices provide great assistance for COVID-19 prevention. However, except for contact tracing, trajectory data are rarely employed in other preventative scenarios against the pandemic. In this article, we try to extend the application of trajectories auto-collected by wireless sensing devices and simulate the epidemic spread in a trajectory data-driven manner. After that, the effects of three nonpharmacological measures are quantified. In contrast to existing studies, additional requirements such as the complex topological networks are needless in our simulation, where the interactions between agents are derived by the intersections of their trajectories. Concretely, the dynamic of virus propagation among individuals is first modeled, and then an agent-based microsimulation environment is built as an artificial system to conduct the epidemic spread simulation. Finally, the trajectories are loaded into the agents as the reliance for their interactions, and the macroscopic changes under different interventions are revealed in a bottom–up way. As a case study, we conduct the simulation based on the trajectories in a real region, in which we find the following. 1) Among the three examined nonpharmacological interventions, community containment is more effective than keeping social distance, which can lower the deaths to nearly 1/9 compared to no action, while travel restrictions play limited roles. 2) There is a strong positive correlation between population densities and mortality. 3) The timing of community containment triggered by confirmed diagnoses is proportional to the number of deaths, thus early containment will significantly decrease mortality.
Zhishuai Li, Gang Xiong 0001, Peijun Ye 0001, Xiaoli Liu 0005, Sasu Tarkoma, Fei-Yue Wang 0001
IEEE Trans. Comput. Soc. Syst.6
2024 The Price is Right? The Economic Value of Sharing Sensors
abstract
We study user's valuations of smartphone sensing resources and the factors mediating them through a systematic auction study with 108 bids from$N=18$participants, two resource use conditions [fixed battery (FB) and variable battery (VB)] and three sensors (camera, microphone, and GPS) with differing energy and privacy costs. We use a second-price sealed-bid reverse auction as this allows us to elicit the participants’ truthful perceived value for sharing resources. We show that most users would be willing to share even highly-privacy intrusive sensors if they are sufficiently compensated. At the FB level, participants placed much lower value for sharing GPS (€13) than camera (€30) or microphone (€32.5). The values people place on sharing access to resources generally reflect four considerations: 1) the perceived value of the sensor type; 2) the value of the data captured by the sensor; 3) the impact of sharing on the device; and 4) personal variations related to sharing motives, personal tendencies, and the broader sharing context. We address the practical impact of our results by presenting two case studies (collaborative sensing and collaborative AI). Finally, we derive design implications for sharing sensing resources on personal devices.
Ngoc Thi Nguyen, Maria Zubair, Agustin Zuniga, Sasu Tarkoma, Pan Hui 0001, Hyowon Lee 0001, Simon T. Perrault, Mostafa H. Ammar, Huber Flores, Petteri Nurmi
IEEE Trans. Comput. Soc. Syst.4
2024 Towards Risk-Averse Edge Computing With Deep Reinforcement Learning
abstract
Recently, artificial intelligence paves the way for the development of smart services for people anytime and anywhere, which poses great challenges on accessing computing resources. Multi-access edge computing complements existing cloud computing infrastructure at the edge of the network, where mobile users can offload computationally intensive tasks of smart applications to edge servers that are in proximity to the users themselves. Existing offloading schemes mainly focus on selecting edge servers for each offloading task with the goal of optimizing the overall average latency. However, the solutions with the optimal overall average latency may be not the most suitable for all offloading tasks. There is still a possibility that offloading leads to an extreme case of ultra-high latency, which is not acceptable for latency-sensitive applications. To address this problem, we therefore introduce modern portfolio theory (MPT) to jointly consider the overall average latency and potential risks in optimal edge server selection. The task offloading problem is regarded as an investment portfolio with the objective of maximizing the ‘return’ while minimizing the risk. Combining MPT with deep reinforcement learning (DRL), we design two proximal policy optimization (PPO)-based task offloading algorithms to jointly optimize these two objectives. The algorithm computes a portfolio for each mobile user that enables the diversification of edge server selection, thereby minimizing the risk and the average latency. Extensive simulation results based on three real-world trace datasets show that our algorithms significantly outperform the state-of-the-art solutions and can reduce the overall average latency and the risk by 59% and 85% at most, respectively.
Dianlei Xu, Xiang Su 0001, Huandong Wang, Sasu Tarkoma, Pan Hui 0001
IEEE Trans. Mob. Comput.4
2024 Adaptive Compression-Aware Split Learning and Inference for Enhanced Network Efficiency
abstract
The growing number of AI-driven applications in mobile devices has led to solutions that integrate deep learning models with the available edge-cloud resources. Due to multiple benefits such as reduction in on-device energy consumption, improved latency, improved network usage, and certain privacy improvements, split learning, where deep learning models are split away from the mobile device and computed in a distributed manner, has become an extensively explored topic. Incorporating compression-aware methods (where learning adapts to compression level of the communicated data) has made split learning even more advantageous. This method could even offer a viable alternative to traditional methods, such as federated learning techniques. In this work, we develop an adaptive compression-aware split learning method (“deprune”) to improve and train deep learning models so that they are much more network-efficient, which would make them ideal to deploy in weaker devices with the help of edge-cloud resources. This method is also extended (“prune”) to very quickly train deep learning models through a transfer learning approach, which tradesoff little accuracy for much more network-efficient inference abilities. We show that the “deprune” method can reduce network usage by 4× when compared with a split-learning approach (that does not use our method) without loss of accuracy, while also improving accuracy over compression-aware split-learning by up to 4 percent. Lastly, we show that the “prune” method can reduce the training time for certain models by up to 6× without affecting the accuracy when compared against a compression-aware split-learning approach.
Akrit Mudvari, Antero Vainio, Iason Ofeidis, Sasu Tarkoma, Leandros Tassiulas
ACM Trans. Internet Techn.4
2024 FedGK: Communication-Efficient Federated Learning through Group-Guided Knowledge Distillation
abstract
Federated learning (FL) empowers a cohort of participating devices to contribute collaboratively to a global neural network model, ensuring that their training data remains private and stored locally. Despite its advantages in computational efficiency and privacy preservation, FL grapples with the challenge of non-IID (not independent and identically distributed) data from diverse clients, leading to discrepancies between local and global models and potential performance degradation. In this article, we propose FedGK, an innovative communication-efficient Group-Guided FL framework designed for heterogeneous data distributions. FedGK employs a localized-guided framework that enables the client to effectively assimilate key knowledge from teachers and peers while minimizing extraneous peer information in FL scenarios. We conduct an in-depth analysis of the dynamic similarities among clients over successive communication rounds and develop a novel clustering approach that accurately groups clients with diverse heterogeneities. We implement FedGK on public datasets with an innovative data transformation pattern called “cluster-shift non-IID”, which mirrors the more prevalent data distributions in real-world settings and could be grouped into clusters with similar data distributions. Extensive experimental results on public datasets demonstrate that the proposed approach FedGK improves accuracy by up to 32.89% and saves up to 53.33% communication cost over state-of-the-art baselines.
Wenjun Zhang 0013, Xiaoli Liu 0005, Sasu Tarkoma
ACM Trans. Internet Techn.3
2024 Behave Differently when Clustering: A Semi-asynchronous Federated Learning Approach for IoT
abstract
The Internet of Things (IoT) has revolutionized the connectivity of diverse sensing devices, generating an enormous volume of data. However, applying machine learning algorithms to sensing devices presents substantial challenges due to resource constraints and privacy concerns. Federated learning (FL) emerges as a promising solution allowing for training models in a distributed manner while preserving data privacy on client devices. We contribute SAFI , a semi-asynchronous FL approach based on clustering to achieve a novel in-cluster synchronous and out-cluster asynchronous FL training mode. Specifically, we propose a three-tier architecture to enable IoT data processing on edge devices and design a clustering selection module to effectively group heterogeneous edge devices based on their processing capacities. The performance of SAFI has been extensively evaluated through experiments conducted on a real-world testbed. As the heterogeneity of edge devices increases, SAFI surpasses the baselines in terms of the convergence time, achieving a speedup of approximately × 3 when the heterogeneity ratio is 7:1. Moreover, SAFI demonstrates favorable performance in non-independent and identically distributed settings and requires lower communication cost compared to FedAsync. Notably, SAFI is the first Java-implemented FL approach and holds significant promise to serve as an efficient FL algorithm in IoT environments.
Boyu Fan, Xiang Su 0001, Sasu Tarkoma, Pan Hui 0001
ACM Trans. Sens. Networks3
2023 Fog Computing for Deep Learning with Pipelines
abstract
In this article, we introduce a fog system design for processing data collected from edge devices, such as mobile, sensor, and extended (mixed, augmented, virtual) reality equipment. Our system enables the network to provide hardware-accelerated processors for resource-intensive computations on data gathered from remote locations, such as 5G and beyond mobile networks. By splitting heavy computations into pipelines, and distributing them among processors in the edge, fog and the cloud, our design benefits from the processing power of the cloud, while utilizing fog devices with a lower network latency. We implement our design, and use it for distributed training and inference with industry-grade deep learning models for computer vision. We deploy our architecture in infrastructure including cloud and edge servers supporting GPU-accelerated computations. We benchmark pipelines in various deployment settings to study the overhead that they introduce. Our contributions are a new design for wide-area data processing, a framework that realizes this design and provides means of developing applications that are optimized in terms of infrastructure and hardware. These contributions are complemented with our benchmark results, which reveal the potential causes of processing overhead.
Antero Vainio, Akrit Mudvari, Diego Kiedanski, Sasu Tarkoma, Leandros Tassiulas
ICFEC4
2023 Poster: Automatic Mass Power Outage Detection in Radio Access Networks
abstract
Mobile devices are expected to be always connected, and this implies that the mobile network is able to quickly identify and address faults that impact service (for example, due to power outages). In this article, we present our approach for automatically detecting mass power outages. Our solution decreases the number of created trouble tickets in two mobile networks by 4.7% and 9.3%.
Milla Lintunen, Gopika Premsankar, Henri Tenhunen, Sasu Tarkoma, Ashwin Rao
MobiSys4
2023 Unmanned Aerial Vehicles for Air Pollution Monitoring: A Survey
abstract
Unmanned Aerial Vehicles (UAVs) equipped with air quality sensors offer a powerful solution for increasing the spatial and temporal resolution of air quality data, searching and detecting emission sources, and monitoring emissions from fixed and mobile sources. Despite the numerous advantages of using UAVs, their use, however, presents several challenges that limit their broader adoption. For example, UAVs require efficient algorithms and components to minimize power consumption, the overall payload used on UAVs needs to be small to ensure optimal portability which poses limitations on the sensors that can be integrated with UAVs, and there is a need for specialized algorithms, e.g., for identifying and locating air pollution sources. Currently, most solutions for UAV-based air quality monitoring focus on specific challenges or demonstrating the potential of using UAVs, and there is a lack of comprehensive overview of the research field and its open challenges. In this paper, we contribute a systematic review of UAV-based air quality monitoring, highlighting and analyzing technical solutions and challenges, and identifying open challenges with the aim of providing a research roadmap for the path forward.
Naser Hossein Motlagh, Pranvera Kortoçi, Xiang Su 0001, Lauri Lovén, Hans Kristian Hoel, Sindre Bjerkestrand Haugsvær, Casper Fabian Gulbrandsen, Petteri Nurmi, Sasu Tarkoma
IEEE Internet Things J.10
2023 Guest Editorial Special Issue on When Blockchain Meets 5G/6G - Enabling Endogenously Secure IoT
abstract
The standardization of the fifth-generation (5G) communications has been completed, and the visioning and planning of the sixth-generation (6G) communications have begun, with an objective of casting the high technical standard of new spectrum, high time and phase synchronization accuracy, and 100% geographical coverage to flexibly and efficiently connect upper trillion-level devices in the future. The transition from 5G to 6G is expected to integrate all operational networks, especially the Internet of Things (IoT), which involves massive heterogeneous devices to interact with our physical world.
Dongxiao Yu, Jian Ren 0001, Sasu Tarkoma, Madhuri Siddula, Falko Dressler
IEEE Internet Things J.4
2023 Context-driven encrypted multimedia traffic classification on mobile devices
abstract
The Internet has been experiencing immense growth in multimedia traffic from mobile devices. The increase in traffic presents many challenges to user-centric networks, network operators, and service providers. Foremost among these challenges is the inability of networks to determine the types of encrypted traffic and thus the level of network service the traffic needs to maintain an acceptable quality of experience. Therefore, end devices are a natural fit for performing traffic classification since end devices have more contextual information about device usage and traffic. This paper proposes a novel approach that classifies multimedia traffic types produced and consumed on mobile devices. The technique relies on a mobile device’s detection of its multimedia context characterized by its utilization of different media input/output (I/O) components, e.g., camera, microphone, and speaker. We develop an algorithm, MediaSense, which senses the states of multiple I/O components and identifies the specific multimedia context of a mobile device in real-time. We demonstrate that MediaSense classifies encrypted multimedia traffic in real-time as accurately as deep learning approaches and with even better generalizability.
Mohammad Ashraful Hoque, Benjamin Finley, Ashwin Rao, Abhishek Kumar 0011, Pan Hui 0001, Mostafa H. Ammar, Sasu Tarkoma
Pervasive Mob. Comput.7
2023 Chirp-Loc: Multi-factor authentication via acoustically-generated location signatures
Prakash Shrestha, Hien Thi Thu Truong, Pupu Toivonen, Nitesh Saxena, Sasu Tarkoma, Petteri Nurmi
Pervasive Mob. Comput.5
2023 Intelligent Air Pollution Sensors Calibration for Extreme Events and Drifts Monitoring
abstract
Air quality low-cost sensors (LCSs) are affordable and can be deployed in massive scale in order to enable high-resolution spatio-temporal air pollution information. However, they often suffer from sensing accuracy, in particular, when they are used for capturing extreme events. We propose an intelligent sensors calibration method that facilitates correcting LCSs measurements accurately and detecting the calibrators’ drift. The proposed calibration method uses Bayesian framework to establish white-box and black-box calibrators. We evaluate the method in a controlled experiment under different types of smoking events. The calibration results show that the method accurately estimates the aerosol mass concentration during the smoking events. We show that black-box calibrators are more accurate than white-box calibrators. However, black-box calibrators may drift easily when a new smoking event occurs, while white-box calibrators remain robust. Therefore, we implement both of the calibrators in parallel to extract both calibrators’ strengths and also enable drifting monitoring for calibration models. We also discuss that our method is implementable for other types of LCSs suffered from sensing accuracy.
Martha Arbayani Zaidan, Naser Hossein Motlagh, Pak Lun Fung, Abedalaziz S. Khalaf, Yutaka Matsumi, Aijun Ding, Sasu Tarkoma, Tuukka Petäjä, Markku Kulmala, Tareq Hussein
IEEE Trans. Ind. Informatics7
2023 Upscaling Fog Computing in Oceans for Underwater Pervasive Data Science Using Low-Cost Micro-Clouds
abstract
Underwater environments are emerging as a new frontier for data science thanks to an increase in deployments of underwater sensor technology. Challenges in operating computing underwater combined with a lack of high-speed communication technology covering most aquatic areas means that there is a significant delay between the collection and analysis of data. This in turn limits the scale and complexity of the applications that can operate based on these data. In this article, we develop underwater fog computing support using low-cost micro-clouds and demonstrate how they can be used to deliver cost-effective support for data-heavy underwater applications. We develop a proof-of-concept micro-cloud prototype and use it to perform extensive benchmarks that evaluate the suitability of underwater micro-clouds for diverse underwater data science scenarios. We conduct rigorous tests in both controlled and field deployments, using river and sea waters. We also address technical challenges in enabling underwater fogs, evaluating the performance of different communication interfaces and demonstrating how accelerometers can be used to detect the likelihood of communication failures and determine which communication interface to use. Our work offers a cost-effective way to increase the scale and complexity of underwater data science applications, and demonstrates how off-the-shelf devices can be adopted for this purpose.
Farooq Dar 0001, Mohan Liyanage, Marko Radeta, Zhigang Yin, Agustin Zuniga, Sokol Kosta, Sasu Tarkoma, Petteri Nurmi, Huber Flores
ACM Trans. Internet Things7
2023 You Are How You Use Apps: User Profiling Based on Spatiotemporal App Usage Behavior
abstract
Mobile apps have become an indispensable part of people’s daily lives. Users determine what apps to use and when and where to use them based on their tastes, interests, and personal demands, depending on their personality traits. This article aims to infer user profiles from their spatiotemporal mobile app usage behavior. Specifically, we first transform mobile app usage records into a heterogeneous graph. On the graph, nodes represent users, apps, locations, and time slots. Edges describe the co-occurrence of entities in usage records. We then develop a multi-relational heterogeneous graph attention network (MRel-HGAN), an end-to-end system for user profiling. MRel-HGAN first adopts a neighbor sampling strategy based on bootstrapping to sample heavily connected neighbors of a fixed size for each node. Next, we design a relational graph convolutional operation and a multi-relational attention operation. Through such modules, MRel-HGAN can generate node embedding by sufficiently leveraging the rich semantic information of the multi-relational structure in the mobile app usage graph. Experimental results on real-world mobile app usage datasets show the effectiveness and superiority of our MRel-HGAN in the user profiling task for attributes of gender and age.
Tong Li 0013, Yong Li 0008, Mingyang Zhang 0004, Sasu Tarkoma, Pan Hui 0001
ACM Trans. Intell. Syst. Technol.4
2023 Learning Representations of Satellite Imagery by Leveraging Point-of-Interests
abstract
Satellite imagery depicts the Earth’s surface remotely and provides comprehensive information for many applications, such as land use monitoring and urban planning. Existing studies on unsupervised representation learning for satellite images only take into account the images’ geographic information, ignoring human activity factors. To bridge this gap, we propose using the Point-of-Interest (POI) data to capture human factors and designing a contrastive learning-based framework to consolidate the representation of satellite imagery with POI information. Besides, we introduce a season-invariant representation learning model on satellite imagery, considering that human factors are mostly unchanging with respect to seasons. An attention model is designed at last to merge the representations from the geographic, seasonal, and POI perspectives adaptively. On the basis of real-world datasets collected from Beijing, 1 we evaluate our method for predicting socioeconomic indicators. The results show that the representation containing POI information outperforms the geographic representation in estimating commercial activity-related indicators. Our proposed attentional framework can estimate the socioeconomic indicators with R 2 of 0.874 and outperforms the baseline methods. Furthermore, we explore the differences in the representations of satellite images with varying socioeconomic statuses. Finally, we investigate the impact of geographic and POI perspective information in the representation learning process, as well as the effect of satellite imagery on various spatial resolutions.
Tong Li 0013, Yanxin Xi, Huandong Wang, Yong Li 0008, Sasu Tarkoma, Pan Hui 0001
ACM Trans. Intell. Syst. Technol.5
2023 Understanding the Long-Term Evolution of Mobile App Usage
abstract
The prevalence of smartphones has promoted the popularity of mobile apps in recent years. Although significant effort has been made to understand mobile app usage, existing studies are based primarily on short-term datasets with a limited time span, e.g., a few months. Therefore, many basic facts about the long-term evolution of mobile app usage are unknown. In this paper, we study how mobile app usage evolves over a long-term period. We first introduce an app usage collection platform named carat, from which we have gathered app usage records of 1,465 users from 2012 to 2017. We then conduct the first study on the long-term evolution processes on a macro-level, i.e., app-category, and micro-level, i.e., individual app. We discover that, on both levels, there is a growth stage enabled by the introduction of new technologies. Then there is a plateau stage caused by high correlations between app categories and a Pareto effect in individual app usage, respectively. Additionally, the evolution of individual app usage undergoes an elimination stage due to fierce intra-category competition. The inter-diversity of app-category and individual app usage exhibits opposing trends: app-category usage assimilates while individual app usage diversifies. Nevertheless, the intra-diversity of both app-category and app usage declines over time. Also, we demonstrate the country barriers of app category usage. We further investigate how different demographics affect the evolutionary processes of app usage. Our study provides useful implications for app developers, market intermediaries, and service providers.
Tong Li 0013, Yali Fan, Yong Li 0008, Sasu Tarkoma, Pan Hui 0001
IEEE Trans. Mob. Comput.4
2022 Predicting Multi-level Socioeconomic Indicators from Structural Urban Imagery
abstract
Understanding economic development and designing government policies requires accurate and timely measurements of socioeconomic activities. In this paper, we show how to leverage city structural information and urban imagery like satellite images and street view images to accurately predict multi-level socioeconomic indicators. Our framework consists of four steps. First, we extract structural information from cities by transforming real-world street networks into city graphs (GeoStruct). Second, we design a contrastive learning-based model to refine urban image features by looking at geographic similarity between images, with images that are geographically close together having similar features (GeoCLR). Third, we propose using street segments as containers to adaptively fuse the features of multi-view urban images, including satellite images and street view images (GeoFuse). Finally, given the city graph with a street segment as a node and a neighborhood area as a subgraph, we jointly model street- and neighborhood-level socioeconomic indicator predictions as node and subgraph classification tasks. The novelty of our method is that we introduce city structure to organize multi-view urban images and model the relationships between socioeconomic indicators at different levels. We evaluate our framework on the basis of real-world datasets collected in multiple cities. Our proposed framework improves performance by over 10% when compared to state-of-the-art baselines in terms of prediction accuracy and recall.
Tong Li 0013, Shiduo Xin, Yanxin Xi, Sasu Tarkoma, Pan Hui 0001, Yong Li 0008
CIKM4
2022 PassWalk: Spatial Authentication Leveraging Lateral Shift and Gaze on Mobile Headsets
abstract
Secure and usable user authentication on mobile headsets is a challenging problem. The miniature-sized touchpad on such devices becomes a hurdle to user interactions that impact usability. However, the most common authentication methods, i.e., the standard QWERTY virtual keyboard or mid-air inputs to enter passwords are highly vulnerable to shoulder surfing attacks. In this paper, we present PassWalk, a keyboard-less authentication system leveraging multi-modal inputs on mobile headsets. PassWalk demonstrates the feasibility of user authentication driven by the user's gaze and lateral shifts (i.e., footsteps) simultaneously. The keyboard-less authentication interface in PassWalk enables users to accomplish highly mobile inputs of graphical passwords, containing digital overlays and physical objects. We conduct an evaluation with 22 recruited participants (15 legitimate users and 7 attackers). Our results show that PassWalk provides high security (only 1.1% observation attacks were successful) with a mean authentication time of 8.028s, which outperforms the commercial method of using the QWERTY virtual keyboard (21.5% successful attacks) and a research prototype LookUnLock (5.5% successful attacks). Additionally, PassWalk entails a significantly smaller workload on the user than the current commercial methods.
Abhishek Kumar 0011, Lik-Hang Lee, Jagmohan Chauhan, Xiang Su 0001, Mohammad Ashraful Hoque, Susanna Pirttikangas, Sasu Tarkoma, Pan Hui 0001
ACM Multimedia7
2022 Context-driven Encrypted Multimedia Traffic Classification on Mobile Devices
abstract
The Internet has been experiencing immense growth in multimedia traffic from mobile devices. The increase in traffic presents many challenges to user-centric networks, network operators, and service providers. Foremost among these challenges is the inability of networks to determine the types of encrypted traffic and thus the level of network service the traffic needs for maintaining an acceptable quality of experience. Therefore, end devices are a natural fit for performing traffic classification since end devices have more contextual information about the device usage and traffic. This paper proposes a novel approach that classifies multimedia traffic types produced and consumed on mobile devices. The technique relies on a mobile device’s detection of its multimedia context characterized by its utilization of different media input/output components, e.g., camera, microphone, and speaker. We develop an algorithm, MediaSense, which senses the states of multiple I/O components and identifies the specific multimedia context of a mobile device in real-time. We demonstrate that MediaSense classifies encrypted multimedia traffic in real-time as accurately as deep learning approaches and with even better generalizability.
Mohammad Ashraful Hoque, Benjamin Finley, Ashwin Rao, Abhishek Kumar 0011, Pan Hui 0001, Mostafa H. Ammar, Sasu Tarkoma
PerCom7
2022 Beyond the First Law of Geography: Learning Representations of Satellite Imagery by Leveraging Point-of-Interests
abstract
Satellite imagery depicts the earth’s surface remotely and provides comprehensive information for many applications, such as land use monitoring and urban planning. Existing studies on unsupervised representation learning for satellite images only take into account the images’ geographic information, ignoring human activity factors. To bridge this gap, we propose using Point-of-Interest (POI) data to capture human factors and design a contrastive learning-based framework to consolidate the representation of satellite imagery with POI information. Also, we design an attention model that merges the representations from the geographic and POI perspectives adaptively. On the basis of real-world datasets collected from Beijing, we evaluate our method for predicting socioeconomic indicators. The results show that the representation containing POI information outperforms the geographic representation in estimating commercial activity-related indicators. Our proposed framework can estimate the socioeconomic indicators with an R2 of 0.874 and outperforms the baseline methods.
Yanxin Xi, Tong Li 0013, Huandong Wang, Yong Li 0008, Sasu Tarkoma, Pan Hui 0001
WWW5
2022 Retrieval of Multiple Atmospheric Environmental Parameters From Images With Deep Learning
abstract
Retrieving atmospheric environmental parameters such as atmospheric horizontal visibility and mass concentration of aerosol particles with a diameter of 2.5 or 10${\mu }\text{m}$or less (PM2.5, PM10, respectively) from digital images provides new tools for horizontal environmental monitoring. In this study, we propose a new end-to-end convolutional neural network (CNN) for the retrieval of multiple atmospheric environmental parameters (RMEPs) from images. In contrast to other retrieval models, RMEP can retrieve a suite of atmospheric environmental parameters including atmospheric horizontal visibility, relative humidity (RH), ambient temperature, PM2.5, and PM10simultaneously from a single image. Experimental results demonstrate that: 1) it is possible to simultaneously retrieve multiple atmospheric environmental parameters; 2) spatial and spectral resolutions of images are not the key factors for the retrieval on the horizontal scale; and 3) RMEP achieves the best overall retrieval performance compared with several classic CNNs such as AlexNet, ResNet-50, and DenseNet-121, and the results are based on experiments on images extracted from webcams located in different continents (test$R^{2}$values are 0.63, 0.72, and 0.82 for atmospheric horizontal visibility, RH, and ambient temperature, respectively). Experimental results show the potential of utilizing webcams to help monitor the environment. Code and more results are available athttps://github.com/cvvsu/RMEP.
Peifeng Su, Yongchun Liu, Sasu Tarkoma, Andrew Rebeiro-Hargrave, Tuukka Petäjä, Markku Kulmala, Petri Pellikka
IEEE Geosci. Remote. Sens. Lett.3
2022 $d$d-Simplexed: Adaptive Delaunay Triangulation for Performance Modeling and Prediction on Big Data Analytics
abstract
Big Data processing systems (e.g., Spark) have a number of resource configuration parameters, such as memory size, CPU allocation, and the number of running nodes. Regular users and even expert administrators struggle to understand the mutual relation between different parameter configurations and the overall performance of the system. In this paper, we address this challenge by proposing a performance prediction framework, called$d$-Simplexed, to build performance models with varied configurable parameters on Spark. We take inspiration from the field of Computational Geometry to construct a$d$-dimensional mesh using Delaunay Triangulation over a selected set of features. From this mesh, we predict execution time for various feature configurations. To minimize the time and resources in building a bootstrap model with a large number of configuration values, we propose an adaptive sampling technique to allow us to collect as few training points as required. Our evaluation on a cluster of computers using WordCount, PageRank, Kmeans, and Join workloads in HiBench benchmarking suites shows that we can achieve less than 5 percent error rate for estimation accuracy by sampling less than 1 percent of data.
Yuxing Chen 0003, Peter Goetsch, Mohammad Ashraful Hoque, Jiaheng Lu, Sasu Tarkoma
IEEE Trans. Big Data5
2022 Agora: A Privacy-Aware Data Marketplace
Vlasis Koutsos, Dimitrios Papadopoulos 0001, Dimitris Chatzopoulos, Sasu Tarkoma, Pan Hui 0001
IEEE Trans. Dependable Secur. Comput.4
2022 Trip Purposes Mining From Mobile Signaling Data
abstract
With the widespread application of mobile phones, it has become possible to study human mobility and travel behaviors based on cellular network data. Contrary to call detail records, the data is triggered by mobile cellular signaling and can provide fine-grained information about users’ daily routines. However, it does not explicitly provide semantic details about traveling traces, e.g., trip purposes. In this paper, we propose a methodological framework to handle large-scale cellular network data and discover the underlying trip purposes in an unsupervised way. We first devise heuristic rules to identify home/work purposes. Then, a flexible latent Dirichlet allocation (LDA) model is presented to discover the activities for remaining trips, in which each trip is depicted by four attributes, i.e. arrival time, age group, stay duration, and the point of interest tag for the destination. Experimental results show that the proposed method can identify diverse trip purposes by explaining their structures over trip attributes and outperform baselines in terms of log-likelihood and perplexity. We also analyze the difference between the automatically discovered trip purposes and those estimated from household census, and the analyzed results demonstrate the feasibility of our proposed method.
Zhishuai Li, Gang Xiong 0001, Zebing Wei, Xiaoli Liu 0005, Sasu Tarkoma, Min Huang 0009, Chuheng Wu
IEEE Trans. Intell. Transp. Syst.7
2022 AICP: Augmented Informative Cooperative Perception
abstract
Connected vehicles, whether equipped with advanced driver-assistance systems or fully autonomous, require human driver supervision and are currently constrained to visual information in their line-of-sight. A cooperative perception system among vehicles increases their situational awareness by extending their perception range. Existing solutions focus on improving perspective transformation and fast information collection. However, such solutions fail to filter out large amounts of less relevant data and thus impose significant network and computation load. Moreover, presenting all this less relevant data can overwhelm the driver and thus actually hinder them. To address such issues, we present Augmented Informative Cooperative Perception (AICP), the first fast-filtering system which optimizes the informativeness of shared data at vehicles to improve the fused presentation. To this end, an informativeness maximization problem is presented for vehicles to select a subset of data to display to their drivers. Specifically, we propose (i) a dedicated system design with custom data structure and lightweight routing protocol for convenient data encapsulation, fast interpretation and transmission, and (ii) a comprehensive problem formulation and efficient fitness-based sorting algorithm to select the most valuable data to display at the application layer. We implement a proof-of-concept prototype of AICP with a bandwidth-hungry, latency-constrained real-life augmented reality application. The prototype adds only 12.6 milliseconds of latency to a current informativeness-unaware system. Next, we test the networking performance of AICP at scale and show that AICP effectively filters out less relevant packets and decreases the channel busy time.
Peng Yuan Zhou, Pranvera Kortoçi, Yui-Pan Yau, Benjamin Finley, Xiujun Wang, Tristan Braud, Lik-Hang Lee, Sasu Tarkoma, Jussi Kangasharju, Pan Hui 0001
IEEE Trans. Intell. Transp. Syst.8
2022 To What Extent We Repeat Ourselves? Discovering Daily Activity Patterns Across Mobile App Usage
abstract
With the prevalence of smartphones, people have left abundant behavior records in cyberspace. Discovering and understanding individuals’ cyber activities can provide useful implications for policymakers, service providers, and app developers. In this paper, we propose a framework to discover daily cyber activity patterns across people's mobile app usage. The framework first segments app usage traces into short time windows and then applies a probabilistic topic model to infer users’ cyber activities in each window. By constructing and exploring the coherence of users’ activity sequences, the framework can identify individuals’ daily patterns. Next, the framework uses a hierarchical clustering algorithm to recognize the common patterns across diverse groups of individuals. We apply the framework on a large-scale and real-world dataset, consisting of 653,092 users with 971,818,946 usage records of 2,000 popular mobile apps. Our analysis shows that people usually follow yesterday's activity patterns, but the patterns tend to deviate as the time-lapse increases. We also discover five common daily cyber activity patterns, including afternoon reading, nightly entertainment, pervasive socializing, commuting, and nightly socializing. Our findings have profound implications on identifying the demographics of users and their lifestyles, habits, service requirements, and further detecting other disrupting trends such as working overtime and addiction to the game and social media.
Tong Li 0013, Yong Li 0008, Mohammad Ashraful Hoque, Tong Xia, Sasu Tarkoma, Pan Hui 0001
IEEE Trans. Mob. Comput.5
2022 IoT vs. Human: A Comparison of Mobility
abstract
Internet of Thing (IoT) devices are rapidly becoming an indispensable part of our life with their increasing deployment in many promising areas, including tele-health, smart city, intelligent agriculture. Understanding the mobility of IoT devices is essential to improve quality of service in IoT applications, such as route planning in logistic management, infrastructure deployment, cellular network update and congestion detection in intelligent traffic. Despite its importance, there are not many results pertaining to the mobility of IoT devices. In this article, we aim to answer three research questions: (i) what are the mobility patterns of IoT device? (ii) what are the differences between IoT device and smartphone mobility patterns? (iii) how the IoT device mobility patterns differ among device types and usage scenarios? We present a comprehensive characterization of IoT device mobility patterns from the perspective of cellular data networks, using a 36-days long signal trace, including 1.5 million IoT devices and 0.425 million smartphones, collected from a nation-wide cellular network in China. We first investigate the basic patterns of IoT devices from two perspectives: temporal and spatial characteristics. Our study finds that IoT device mobility exhibits significantly different patterns compared with smartphones in multiple aspects. For instance, IoT devices move more frequently and have larger radius of gyration. Then we explore the essential mobility of IoT devices by utilizing two models that reveal the nature of human mobility, i.e., exploration and preferential return (EPR) model and entropy based predictability model. We find that IoT devices, with few exceptions, behave totally different from human, and we further derive a new formulation to describe their movement. We also find the gap mobility predictability and predictability limit between IoT and human is not as big as people expected.
Dianlei Xu, Huandong Wang, Yong Li 0008, Sasu Tarkoma, Depeng Jin, Pan Hui 0001
IEEE Trans. Mob. Comput.4
2022 Quality of Monitoring for Cellular Networks
abstract
5G networks and beyond introduce a larger number of Network Elements (NEs) and functions than former cellular generations. The increase in NEs will, thus, result in significantly increasing the Management-Plane (M-Plane) data collected from the NEs. Therefore, the conventional centralized Network Management Systems (NMSs) will face fundamental challenges in processing the M-Plane data. In this paper, we present the concept of Quality of Monitoring (QoM) as a solution, which is able to reduce the M-Plane data already at the NEs. First, QoM aggregates the raw M-Plane data into Key Performance Indicators (KPIs). To these KPIs, the QoM applies a data-driven algorithm to define information loss limits for QoM classes specific for each KPI time series. Then, the QoM applies the classes for compressing the KPI data utilizing a lossy-compression method, which is a derivative of the Piece-Wise Constant Approximation (PWCA) algorithm. To evaluate the performance of the QoM solution, we use M-Plane raw data from a live LTE network and calculate four KPIs, while each KPI has different statistical characteristics. We also define three QoM classes namedExact,Optimized, andSharp. For all KPIs, the classOptimizedhas a higher compression rate than the classExact, while the classSharphas the highest compression rate. Assuming that, for example, NEs of a network produce 280 MB of raw data containing information that needs to be transferred to the network operations center; we use KPIs to represent the information contents of the data, and QoM solution to transfer the data over the network. As a result, the QoM solution achieves an estimated 95% compression gain from the raw data in transfer.
Naser Hossein Motlagh, Shubham Kapoor, Rola Alhalaseh, Sasu Tarkoma, Kimmo Hätönen
IEEE Trans. Netw. Serv. Manag.4
2022 Evaluation of Low-cost Air Quality Sensor Calibration Models
abstract
We contribute a novel model evaluation technique that divides available measurements into training and testing sets in a way that adheres to the requirements imposed on professional monitoring stations. We perform extensive and systematic experiments with a wide range of state-of-the-art calibration models to demonstrate that our approach provides accurate insights about the performance of calibration models in real-world deployments, while at the same time highlighting issues with evaluation techniques used in previous works. Among others, our results show that although trained and tested in the same location, calibration errors can exhibit deviation up to 116% depending on the evaluation protocol that is being adopted. We also demonstrate that models trained with continuous data can suffer up to 76% greater error when tested with data coming from diverse environmental conditions. In contrast, when models are trained and tested with our method, the variability of errors is significantly reduced and the robustness of calibration models is significantly improved. The overall performance improvements depend on pollutant concentration, ranging from 10% for low concentrations to 90% for high concentrations that represent conditions that are most dangerous for human health.
Kasimir Aula, Eemil Lagerspetz, Petteri Nurmi, Sasu Tarkoma
ACM Trans. Sens. Networks4
2021 Context-Aware Augmented Reality with 5G Edge
abstract
Augmented Reality (AR) provides immersive user experiences by overlaying digital information on physical environments. Context-awareness is crucial for delivering relevant augmentations that best suit users' requirements and their en-vironments. In this article, we combine context-aware reasoning with emerging AR applications to provide the most relevant infor-mation according to user and environment contexts. To support the best possible quality of experience, 5G edge computing enables the distribution of computation-intensive AR tasks to edge servers through 5G networks. We develop ConAR, a context-aware head-mounted display AR system that is deployed on the edge and cloud leveraging both environmental sensors and user profile context for navigation. ConAR is composed of a HoloLens application and a paired mobile client, which contains a context model for air quality forecasting, and rendering recommendations on holograms through a HoloLens 2 device. We evaluate our system performance by deploying our proposed air quality prediction algorithm on the edge and cloud while communicating to them using 5G and LTE connections. We measure network quality metrics and find the deployment on the edge with 5G connections significantly outperforms alternative solutions. Our results demonstrate that the 5G edge computing is suitable for supporting latency-sensitive analysis tasks for context-aware AR.
Jacky Cao, Xiaoli Liu 0005, Xiang Su 0001, Sasu Tarkoma, Pan Hui 0001
GLOBECOM4
2021 Demographics of mobile app usage: long-term analysis of mobile app usage
Zhen Tu, Hancheng Cao, Eemil Lagerspetz, Yali Fan, Huber Flores, Sasu Tarkoma, Petteri Nurmi, Yong Li 0008
CCF Trans. Pervasive Comput. Interact.6
2021 The Impact of Covid-19 on Smartphone Usage
abstract
The outbreak of Covid-19 changed the world as well as human behavior. In this article, we study the impact of Covid-19 on smartphone usage. We gather smartphone usage records from a global data collection platform called Carat, including the usage of mobile users in North America from November 2019 to April 2020. We then conduct the first study on the differences in smartphone usage across the outbreak of Covid-19. We discover that Covid-19 leads to a decrease in users' smartphone engagement and network switches, but an increase in WiFi usage. Also, its outbreak causes new typical diurnal patterns of both memory usage and WiFi usage. Additionally, we investigate the correlations between smartphone usage and daily confirmed cases of Covid-19. The results reveal that memory usage, WiFi usage, and network switches of smartphones have significant correlations, whose absolute values of Pearson coefficients are greater than 0.8. Moreover, smartphone usage behavior has the strongest correlation with the Covid-19 cases occurring after it, which exhibits the potential of inferring outbreak status. By conducting extensive experiments, we demonstrate that for the inference of outbreak stages, both Macro-F1 and Micro-F1 can achieve over 0.8. Our findings explore the values of smartphone usage data for fighting against the epidemic.
Tong Li 0013, Mingyang Zhang 0004, Yong Li 0008, Eemil Lagerspetz, Sasu Tarkoma, Pan Hui 0001
IEEE Internet Things J.5
2021 Aggregate Cyber-Risk Management in the IoT Age Cautionary Statistics for (Re)Insurers and Likes
abstract
IoT-driven smart societies are modern service-networked ecosystems, whose proper functioning is hugely based on the success of supply chain relationships. Robust security is still a big challenge in such ecosystems, catalyzed primarily by naive cyber-security practices (e.g., setting default IoT device passwords) on behalf of the ecosystem managers, i.e., users and organizations. This has recently led to some catastrophic malware-driven DDoS and ransomware attacks (e.g., the Mirai and WannaCry attacks). Consequently, markets for commercial third-party cyber-risk management (CRM) services (e.g., cyber-insurance) are steadily but sluggishly gaining traction with the rapid increase of IoT deployment in society, and provides a channel for ecosystem managers to transfer residual cyber-risk post attack events. Current empirical studies have shown that such residual cyber-risks affecting smart societies are often heavy-tailed in nature and exhibit tail dependencies. This is both, a major concern for a profit-minded CRM firm that might normally need to cover multiple such dependent cyber-risks from different sectors (e.g., manufacturing and energy) in a service-networked ecosystem, and a good intuition behind the sluggish market growth of CRM products. In this article, we provide: 1) a rigorous general theory to elicit conditions on (tail-dependent) heavy-tailed cyber-risk distributions under which a risk management firm might find it (non)sustainable to provide aggregate cyber-risk coverage services for smart societies and 2) a real-data-driven numerical study to validate claims made in theory assuming boundedly rational cyber-risk managers, alongside providing ideas to boost markets that aggregate dependent cyber-risks with heavy-tails. To the best of our knowledge, this is the only complete general theory till date on the feasibility of aggregate CRM.
Ranjan Pal, Ziyuan Huang 0004, Xinlong Yin, Sergey V. Lototsky, Swades De, Sasu Tarkoma, Mingyan Liu, Jon Crowcroft, Nishanth Sastry
IEEE Internet Things J.6
2021 Corrections to Aggregate Cyber-Risk Management in the IoT Age: Cautionary Statistics for (Re)Insurers and Likes
abstract
As authors of our recently accepted article:Aggregate Cyber-Risk Management in the IoT Age: Cautionary Statistics for (Re)Insurers and Likes, published in the IEEE IoT Journal, we regret that we have found a few errors in the numerical evaluation setup of the works in[1]and[2]that we had borrowed for our accepted paper. In this correction statement, we describe the errors in detail, correct it, and present our revised results with a renewed experimental setup, hoping it to replace the existing incorrect numerical results in the accepted paper. We apologize for the inconvenience caused to the reader. We emphasize that the numerical evaluation section does not in any way hamper the theoretical contributions in this article, and was initially only meant to provide some empirical evidence for whether the theory proposed in this article generalizes to behavioral settings introduced in[2].
Ranjan Pal, Ziyuan Huang 0004, Xinlong Yin, Sergey V. Lototsky, Swades De, Sasu Tarkoma, Mingyan Liu, Jon Crowcroft, Nishanth Sastry
IEEE Internet Things J.6
2021 Edge Intelligence: Empowering Intelligence to the Edge of Network
abstract
Edge intelligence refers to a set of connected systems and devices for data collection, caching, processing, and analysis proximity to where data are captured based on artificial intelligence. Edge intelligence aims at enhancing data processing and protects the privacy and security of the data and users. Although recently emerged, spanning the period from 2011 to now, this field of research has shown explosive growth over the past five years. In this article, we present a thorough and comprehensive survey of the literature surrounding edge intelligence. We first identify four fundamental components of edge intelligence, i.e., edge caching, edge training, edge inference, and edge offloading based on theoretical and practical results pertaining to proposed and deployed systems. We then aim for a systematic classification of the state of the solutions by examining research results and observations for each of the four components and present a taxonomy that includes practical problems, adopted techniques, and application goals. For each category, we elaborate, compare, and analyze the literature from the perspectives of adopted techniques, objectives, performance, advantages and drawbacks, and so on. This article provides a comprehensive survey of edge intelligence and its application areas. In addition, we summarize the development of the emerging research fields and the current state of the art and discuss the important open issues and possible theoretical and technical directions.
Dianlei Xu, Tong Li 0013, Yong Li 0008, Xiang Su 0004, Sasu Tarkoma, Tao Jiang 0002, Jon Crowcroft, Pan Hui 0001
Proc. IEEE5
2021 Low-Cost Outdoor Air Quality Monitoring and Sensor Calibration: A Survey and Critical Analysis
abstract
The significance of air pollution and the problems associated with it are fueling deployments of air quality monitoring stations worldwide. The most common approach for air quality monitoring is to rely on environmental monitoring stations, which unfortunately are very expensive both to acquire and to maintain. Hence, environmental monitoring stations are typically sparsely deployed, resulting in limited spatial resolution for measurements. Recently, low-cost air quality sensors have emerged as an alternative that can improve the granularity of monitoring. The use of low-cost air quality sensors, however, presents several challenges: They suffer from cross-sensitivities between different ambient pollutants; they can be affected by external factors, such as traffic, weather changes, and human behavior; and their accuracy degrades over time. Periodic re-calibration can improve the accuracy of low-cost sensors, particularly with machine-learning-based calibration, which has shown great promise due to its capability to calibrate sensors in-field. In this article, we survey the rapidly growing research landscape of low-cost sensor technologies for air quality monitoring and their calibration using machine learning techniques. We also identify open research challenges and present directions for future research.
Francesco Concas, Julien Mineraud, Eemil Lagerspetz, Samu Varjonen, Xiaoli Liu 0005, Kai Puolamäki, Petteri Nurmi, Sasu Tarkoma
ACM Trans. Sens. Networks8
2020 Agora: A Privacy-aware Data Marketplace
abstract
We propose Agora, the first privacy-aware data marketplace that enables parties to get compensated for contributing data, without relying on a trusted third party. We leverage cryptographic techniques to achieve three security properties: (i) data privacy-raw data remain private except for a function output, (ii) output verifiability-the output is proven to be correct, and (iii) atomicity of payments-parties cannot avoid paying for provided services. Agora is designed as a decentralized blockchain application via smart contracts. We implement a prototype on Ethereum and evaluate its performance in terms of computation overhead and monetary cost.
Vlasis Koutsos, Dimitrios Papadopoulos 0001, Dimitris Chatzopoulos, Sasu Tarkoma, Pan Hui 0001
ICDCS4
2020 Poster: User-space Networking Libraries & Control Plane Negotiations for Seamless Multi-connectivity
abstract
Offering seamless connectivity to devices capable of simultaneously using multiple communication interfaces continues to be a hard problem. This problem is important for edge computing because edge services may be available only on a subset of networks to which the device is capable of connecting to. We argue that various aspects of this problem can be addressed by leveraging the current trends of using user space libraries for networking, and allowing control plane negotiations between user devices and networks.
Seppo Hätönen, Ashwin Rao, Sasu Tarkoma
SEC3
2020 The bits of silence: redundant traffic in VoIP
abstract
Human conversation is characterized by brief pauses and so-called turn-taking behavior between the speakers. In the context of VoIP, this means that there are frequent periods where the microphone captures only background noise - or even silence whenever the microphone is muted. The bits transmitted from such silence periods introduce overhead in terms of data usage, energy consumption, and network infrastructure costs. In this paper, we contribute by shedding light on these costs for VoIP applications. We systematically measure the performance of six popular mobile VoIP applications with controlled human conversation and acoustic setup. Our analysis demonstrates that significant savings can indeed be achieved - with the best performing silence suppression technique being effective on 75% of silent pauses in the conversation in a quiet place. This results in 2-5 times data savings, and 50-90% lower energy consumption compared to the next best alternative. Even then, the effectiveness of silence suppression can be sensitive to the amount of background noise, underlying speech codec, and the device being used. The codec characteristics and performance do not depend on the network type. However, silence suppression makes VoIP traffic network friendly as much as VoLTE traffic. Our results provide new insights into VoIP performance and offer a motivation for further enhancements to a wide variety of voice assisted applications, such as home assistants and other IoT devices.
Mohammad Ashraful Hoque, Petteri Nurmi, Matti Siekkinen, Pan Hui 0001, Sasu Tarkoma
MMSys5
2020 Sensing multimedia contexts on mobile devices
abstract
We use various multimedia applications on smart devices to consume multimedia content, to communicate with our peers, and to broadcast our events live. This paper investigates the utilization of different media input/output devices, e.g., camera, microphone, and speaker, by different types of multimedia applications, and introduces the notion of multimedia context. Our measurements lead to a sensing algorithm called MediaSense, which senses the states of multiple I/O devices and identifies eleven multimedia contexts of a mobile device in real time. The algorithm distinguishes stored content playback from streaming, live broadcasting from local recording, and conversational multimedia sessions from GSM/VoLTE calls on mobile devices.
Mohammad Ashraful Hoque, Ashwin Rao, Abhishek Kumar 0011, Mostafa H. Ammar, Pan Hui 0001, Sasu Tarkoma
NOSSDAV6
2020 COSINE: Collaborator Selector for Cooperative Multi-Device Sensing and Computing
abstract
Pervasive availability of programmable smart de-vices is giving rise to sensing and computing scenarios that involve collaboration between multiple devices. Maximizing the benefits of collaboration requires careful selection of devices with whom to collaborate as otherwise collaboration may be interrupted prematurely or be sub-optimal for the characteristics of the task at hand. Existing research on collaborative scenarios has mostly focused on providing mechanisms that can establish and harness collaboration, without considering how to maximally benefit from it. In this paper, we contribute by developing COSINE as a novel approach for selecting collaborators in multi-device computing scenarios. COSINE identifies and recommends collaborators based on a novel information theoretic measure based on Markov trajectory entropy. Rigorous experimental benchmarks carried out using a large-scale dataset of device-to-device encounters demonstrate that COSINE can significantly improve collaboration benefits compared to current state-of-the-art solutions, increasing expected duration of collaboration and reducing variability of collaborations.
Huber Flores, Agustin Zuniga, Farbod Faghihi, Samuli Hemminki, Sasu Tarkoma, Pan Hui 0001, Petteri Nurmi
PerCom6
2020 BONIK: A Blockchain Empowered Chatbot for Financial Transactions
abstract
A Chatbot is a popular platform to enable users to interact with a software or website to gather information or execute actions in an automated fashion. In recent years, chatbots are being used for executing financial transactions, however, there are a number of security issues, such as secure authentication, data integrity, system availability and transparency, that must be carefully handled for their wide-scale adoption. Recently, the blockchain technology, with a number of security advantages, has emerged as one of the foundational technologies with the potential to disrupt a number of application domains, particularly in the financial sector. In this paper, we forward the idea of integrating a chatbot with blockchain technology in the view to improve the security issues in financial chatbots. More specifically, we present BONIK, a blockchain empowered chatbot for financial transactions, and discuss its architecture and design choices. Furthermore, we explore the developed Proof-of-Concept (PoC), evaluate its performance, analyse how different security and privacy issues are mitigated using BONIK.
Md. Saiful Islam Bhuiyan, Abdur Razzak, Md Sadek Ferdous, Mohammad Jabed Morshed Chowdhury, Mohammad Ashraful Hoque, Sasu Tarkoma
TrustCom6
2020 "What Apps Did You Use?": Understanding the Long-term Evolution of Mobile App Usage
abstract
The prevalence of smartphones has promoted the popularity of mobile apps in recent years. Although significant effort has been made to understand mobile app usage, existing studies are based primarily on short-term datasets with limited time span, e.g., a few months. Therefore, many basic facts about the long-term evolution of mobile app usage are unknown. In this paper, we study how mobile app usage evolves over a long-term period. We first introduce an app usage collection platform named carat, from which we have gathered app usage records of 1,465 users from 2012 to 2017. We then conduct the first study on the long-term evolution processes on a macro-level, i.e., app-category, and micro-level, i.e., individual app. We discover that, on both levels, there is a growth stage enabled by the introduction of new technologies. Then there is a plateau stage caused by high correlations between app categories and a pareto effect in individual app usage, respectively. Additionally, the evolution of individual app usage undergoes an elimination stage due to fierce intra-category competition. Nevertheless, the diverseness of app-category and individual app usage exhibit opposing trends: app-category usage assimilates while individual app usage diversifies. Our study provides useful implications for app developers, market intermediaries, and service providers.
Tong Li 0013, Mingyang Zhang 0004, Hancheng Cao, Yong Li 0008, Sasu Tarkoma, Pan Hui 0001
WWW5
2020 Human Data Model: Improving Programmability of Health and Well-Being Data for Enhanced Perception and Interaction
abstract
Today, an increasing number of systems produce, process, and store personal and intimate data. Such data has plenty of potential for entirely new types of software applications, as well as for improving old applications, particularly in the domain of smart healthcare. However, utilizing this data, especially when it is continuously generated by sensors and other devices, with the current approaches is complex—data is often using proprietary formats and storage, and mixing and matching data of different origin is not easy. Furthermore, many of the systems are such that they should stimulate interactions with humans, which further complicates the systems. In this article, we introduce the Human Data Model—a new tool and a programming model for programmers and end users with scripting skills that help combine data from various sources, perform computations, and develop and schedule computer-human interactions. Written in JavaScript, the software implementing the model can be run on almost any computer either inside the browser or using Node.js. Its source code can be freely downloaded from GitHub, and the implementation can be used with the existing IoT platforms. As a whole, the work is inspired by several interviews with professionals, and an online survey among healthcare and education professionals, where the results show that the interviewed subjects almost entirely lack ideas on how to benefit the ever-increasing amount of data measured of the humans. We believe that this is because of the missing support for programming models for accessing and handling the data, which can be satisfied with the Human Data Model.
Niko Mäkitalo, Daniel Flores-Martin, Huber Flores, Eemil Lagerspetz, François Christophe, Petri Ihantola, Masiar Babazadeh, Pan Hui 0001, Juan Manuel Murillo, Sasu Tarkoma, Tommi Mikkonen
ACM Trans. Comput. Heal.10
2020 On Reconfiguring 5G Network Slices
abstract
The virtual resources of 5G networks are expected to scale and support migration to other locations within the substrate. In this context, a configuration for 5G network slices details the instantaneous mapping of the virtual resources across all slices on the substrate, and a feasible configuration satisfies the Service-Level Objectives (SLOs) without overloading the substrate. Reconfiguring a network from a given source configuration to the desired target configuration involves identifying an ordered sequence of feasible configurations from the source to the target. The proposed solutions for finding such a sequence are optimized for data centers and cannot be used as-is for reconfiguring 5G network slices. We present Matryoshka, our divide-and-conquer approach for finding a sequence of feasible configurations that can be used to reconfigure 5G network slices. Unlike previous approaches, Matryoshka also considers the bandwidth and latency constraints between the network functions of network slices. Evaluating Matryoshka required a dataset of pairs of source and target configurations. Because such a dataset is currently unavailable, we analyze proof of concept roll-outs, trends in standardization bodies, and research sources to compile an input dataset. On using Matryoshka on our dataset, we observe that it yields close-to-optimal reconfiguration sequences 10X faster than existing approaches.
Matteo Pozza, Patrick K. Nicholson, Diego Lugones, Ashwin Rao, Hannu Flinck, Sasu Tarkoma
IEEE J. Sel. Areas Commun.6
2020 Multiple Set Matching with Bloom Matrix and Bloom Vector
abstract
Bloom Filter is a space-efficient probabilistic data structure for checking the membership of elements in a set. Given multiple sets, a standard Bloom Filter is not sufficient when looking for the items to which an element or a set of input elements belong. An example case is searching for documents with keywords in a large text corpus, which is essentially a multiple set matching problem where the input is single or multiple keywords, and the result is a set of possible candidate documents. This article solves the multiple set matching problem by proposing two efficient Bloom Multifilters called Bloom Matrix and Bloom Vector, which generalize the standard Bloom Filter. Both structures are space-efficient and answer queries with a set of identifiers for multiple set matching problems. The space efficiency can be optimized according to the distribution of labels among multiple sets: Uniform and Zipf. Bloom Vector efficiently exploits the Zipf distribution of data for further space reduction. Indeed, both structures are much more space-efficient compared with the state-of-the-art, Bloofi. The results also highlight that a L ookup operation on Bloom Matrix is significantly faster than on Bloom Vector and Bloofi.
Francesco Concas, Pengfei Xu 0004, Mohammad Ashraful Hoque, Jiaheng Lu, Sasu Tarkoma
ACM Trans. Knowl. Discov. Data5
2020 IoT-KEEPER: Detecting Malicious IoT Network Activity Using Online Traffic Analysis at the Edge
abstract
IoT devices are notoriously vulnerable even to trivial attacks and can be easily compromised. In addition, resource constraints and heterogeneity of IoT devices make it impractical to secure IoT installations using traditional endpoint and network security solutions. To address this problem, we present IoT-Keeper, a lightweight system which secures the communication of IoT. IoT-Keeper uses our proposed anomaly detection technique to perform traffic analysis at edge gateways. It uses a combination of fuzzy C-means clustering and fuzzy interpolation scheme to analyze network traffic and detect malicious network activity. Once malicious activity is detected, IoT-Keeper automatically enforces network access restrictions against IoT device generating this activity, and prevents it from attacking other devices or services. We have evaluated IoT-Keeper using a comprehensive dataset, collected from a real-world testbed, containing popular IoT devices. Using this dataset, our proposed technique achieved high accuracy (≈0.98) and low false positive rate (≈0.02) for detecting malicious network activity. Our evaluation also shows that IoT-Keeper has low resource footprint, and it can detect and mitigate various network attacks-without requiring explicit attack signatures or sophisticated hardware.
Ibbad Hafeez, Markku Antikainen, Aaron Yi Ding, Sasu Tarkoma
IEEE Trans. Netw. Serv. Manag.4
2019 The Impact of Thread-Per-Core Architecture on Application Tail Latency
abstract
The response time of an online service depends on the tail latency of a few of the applications it invokes in parallel to satisfy the requests. The individual applications are composed of one or more threads to fully utilize the available CPU cores, but this approach can incur serious overheads. The thread-per-core architecture has emerged to reduce these overheads, but it also has its challenges from thread synchronization and OS interfaces. Applications can mitigate both issues with different techniques, but their impact on application tail latency is an open question. We measure the impact of thread-per-core architecture on application tail latency by implementing a key-value store that uses application-level partitioning, and inter-thread messaging and compare its tail latency to Memcached which uses a traditional key-value store design. We show in an experimental evaluation that our approach reduces tail latency by up to 71 % compared to baseline Memcached running on commodity hardware and Linux. However, we observe that the thread-per-core approach is held back by request steering and OS interfaces, and it could be further improved with NIC hardware offload.
Pekka Enberg, Ashwin Rao, Sasu Tarkoma
ANCS3
2019 Cost-effective Resource Provisioning for Spark Workloads
abstract
Spark is one of the prevalent big data analytical platforms. Configuring proper resource provision for Spark jobs is challenging but essential for organizations to save time, achieve high resource utilization, and remain cost-effective. In this paper, we study the challenge of determining the proper parameter values that meet the performance requirements of workloads while minimizing both resource cost and resource utilization time. We propose a simulation-based cost model to predict the performance of jobs accurately. We achieve low-cost training by taking advantage of simulation framework, i.e., Monte Carlo (MC) simulation, which uses a small amount of data and resources to make a reliable prediction for larger datasets and clusters. The salient feature of our method is that it allows us to invest low training cost while obtaining an accurate prediction. Through experiments with six benchmark workloads, we demonstrate that the cost model yields less than 7% error on average prediction accuracy and the recommendation achieves up to 5x resource cost saving.
Yuxing Chen 0003, Jiaheng Lu, Mohammad Ashraful Hoque, Sasu Tarkoma
CIKM5
2019 I/O Is Faster Than the CPU: Let's Partition Resources and Eliminate (Most) OS Abstractions
abstract
I/O is getting faster in servers that have fast programmable NICs and non-volatile main memory operating close to the speed of DRAM, but single-threaded CPU speeds have stagnated. Applications cannot take advantage of modern hardware capabilities when using interfaces built around abstractions that assume I/O to be slow. We therefore propose a structure for an OS called parakernel, which eliminates most OS abstractions and provides interfaces for applications to leverage the full potential of the underlying hardware. The parakernel facilitates application-level parallelism by securely partitioning the resources and multiplexing only those resources that are not partitioned.
Pekka Enberg, Ashwin Rao, Sasu Tarkoma
HotOS3
2019 Enhancing Indoor IoT Communication with Visible Light and Ultrasound
abstract
The number of deployed Internet of Things (IoT) devices is steadily increasing to manage and interact with community assets of smart cities, such as transportation systems and power plants. This may lead to degraded network performance due to the growing amount of network traffic and connections generated by various IoT devices. To tackle these issues, one promising direction is to leverage the physical proximity of communicating devices and inter-device communication to achieve low latency, bandwidth efficiency, and resilient services. In this work, we aim at enhancing the performance of indoor IoT communication (e.g., smart homes, SOHO) by taking advantage of emerging technologies such as visible light and ultrasound. This approach increases the network capacity, robustness of network connections across IoT devices, and provides efficient means to enable distance-bounding services. We have developed communication modules using off-the-shelf components for visible light and ultrasound and evaluate their network performance and energy consumption. In addition, we show the efficacy of our communication modules by applying them in a practical indoor IoT scenario to realize secure IoT group communication.
Michael Haus, Aaron Yi Ding, Qing Wang 0007, Juhani Toivonen, Leonardo Tonetto, Sasu Tarkoma, Jörg Ott
ICC6
2019 MegaSense: Feasibility of Low-Cost Sensors for Pollution Hot-spot Detection
abstract
Air pollution is a major problem in urban areas, where high population density is accompanied with excess anthropomorphic emissions impacting the environment and increasing health effects. Highly accurate air quality monitoring stations have been used to monitor the severity of the problem and warn citizens. However, air quality can vary sharply even within the same city block, and pollution exposure can vary even 30% between individuals living in the same residence. Therefore, a dense deployment of air quality sensors is needed to detect these variations, and protect citizens from overexposure. Low-cost air quality sensors make it possible to densely instrument a city and detect hot spots as they happen. However, thus far limited information exists on their accuracy and practicability. In this paper, we conduct a 44-day measurement campaign to assess performance of low-cost air quality monitors under different environmental conditions. As practical use case, we consider pollution hot spot detection. Our results show that the mean error of low-cost sensors is small, but the variation in error is significantly larger than with reference sensors. We also show that the accuracy is sufficient for applications relying on variations in air quality index values, such as hot spot detection.
Eemil Lagerspetz, Sasu Tarkoma, Tareq Hussein, Naser Hossein Motlagh, Martha Arbayani Zaidan, Pak Lun Fung, Julien Mineraud, Samu Varjonen, Matti Siekkinen, Petteri Nurmi, Yutaka Matsumi
INDIN2
2019 Indoor Air Quality Monitoring Using Infrastructure-Based Motion Detectors
abstract
Poor indoor air quality is a significant burden to society that can cause health issues and decrease productivity. According to research, indoor air quality is intrinsically linked with human activity and mobility. Indeed, mobility is directly linked with transfer of small particles (e.g. PM2.5) and extent of activity affects production of CO2. Currently, however, estimation of indoor quality is difficult, requiring deployment of highly specialized sensing devices which need to be carefully placed and maintained. In this paper, we contribute by examining the suitability of infrastructure-based motion detectors for indoor air quality estimation. Such sensors are increasingly being deployed into smart environments, e.g., to control lighting and ventilation for energy management purposes. Being able to take advantage of these sensors would thus provide a cost-effective solution for indoor quality monitoring without need for deploying additional sensors. We perform a feasibility study considering measurements collected from a smart office environment having a dense deployment of motion detectors and correlating measurements obtained from motion detectors against air quality values. We consider two main pollutants, PM2.5and CO2, and demonstrate that there indeed is a connection between extent of movement and PM2.5concentration. However, for CO2, no relationship can be established, mostly due to difficulties in separating between people passing by and those residing long-term in the environment.
Naser Hossein Motlagh, Petteri Nurmi, Sasu Tarkoma, Martha Arbayani Zaidan, Eemil Lagerspetz, Samu Varjonen, Juhani Toivonen, Julien Mineraud, Andrew Rebeiro-Hargrave, Matti Siekkinen, Tareq Hussein
INDIN3
2019 Composing 5G Network Slices by Co-locating VNFs in μslices
abstract
5G networks leverage network slices for serving use cases with different and potentially conflicting requirements. Current approaches for composing network slices largely overlook the presence of multiple ingress and egress nodes in each network slice. This results in inefficient resource usage even when co-locating the Virtual Network Functions (VNFs) of a slice. The network is thus quickly saturated, and no more use cases can be served. We address this issue by introducing μslices, where each μslice is tied to a specific pair of ingress and egress nodes of a slice. Then, we optimize the bandwidth consumed by each μslice by co-locating its VNF instances that communicate the most. We observed that enabling μslicing was able to save about two times more control and data plane traffic and it also resulted in a more even distribution of computational and link load.
Matteo Pozza, Akanksha Patel, Ashwin Rao, Hannu Flinck, Sasu Tarkoma
Networking5
2019 DoubleEcho: Mitigating Context-Manipulation Attacks in Copresence Verification
abstract
Copresence verification based on context can improve usability and strengthen security of many authentication and access control systems. By sensing and comparing their surroundings, two or more devices can tell whether they are copresent and use this information to make access control decisions. To the best of our knowledge, all context-based copresence verification mechanisms to date are susceptible to context-manipulation attacks. In such attacks, a distributed adversary replicates the same context at the (different) locations of the victim devices, and induces them to believe that they are copresent. In this paper we propose DoubleEcho, a context-based copresence verification technique that leverages acoustic Room Impulse Response (RIR) to mitigate context-manipulation attacks. In DoubleEcho, one device emits a wide-band audible chirp and all participating devices record reflections of the chirp from the surrounding environment. Since RIR is, by its very nature, dependent on the physical surroundings, it constitutes a unique location signature that is hard for an adversary to replicate. We evaluate DoubleEcho by collecting RIR data with various mobile devices and in a range of different locations. We show that DoubleEcho mitigates context-manipulation attacks whereas all other approaches to date are entirely vulnerable to such attacks. DoubleEcho detects copresence (or lack thereof) in roughly 2 seconds and works on commodity devices.
Hien Thi Thu Truong, Juhani Toivonen, Thien Duc Nguyen, Claudio Soriente, Sasu Tarkoma, N. Asokan
PerCom5
2019 Tortoise or Hare? Quantifying the Effects of Performance on Mobile App Retention
abstract
We contribute by quantifying the effect of network latency and battery consumption on mobile app performance and retention, i.e., user's decisions to continue or stop using apps. We perform our analysis by fusing two large-scale crowdsensed datasets collected by piggybacking on information captured by mobile apps. We find that app performance has an impact in its retention rate. Our results demonstrate that high energy consumption and high latency decrease the likelihood of retaining an app. Conversely, we show that reducing latency or energy consumption does not guarantee higher likelihood of retention as long as they are within reasonable standards of performance. However, we also demonstrate that what is considered reasonable depends on what users have been accustomed to, with device and network characteristics, and app category playing a role. As our second contribution, we develop a model for predicting retention based on performance metrics. We demonstrate the benefits of our model through empirical benchmarks which show that our model not only predicts retention accurately, but generalizes well across application categories, locations and other factors moderating the effect of performance.
Agustin Zuniga, Huber Flores, Eemil Lagerspetz, Petteri Nurmi, Sasu Tarkoma, Pan Hui 0001, Jukka Manner
WWW5
2019 Exploiting Usage to Predict Instantaneous App Popularity: Trend Filters and Retention Rates
abstract
Popularity of mobile apps is traditionally measured by metrics such as the number of downloads, installations, or user ratings. A problem with these measures is that they reflect usage only indirectly. Indeed, retention rates, i.e., the number of days users continue to interact with an installed app, have been suggested to predict successful app lifecycles. We conduct the first independent and large-scale study of retention rates and usage trends on a dataset of app-usage data from a community of 339,842 users and more than 213,667 apps. Our analysis shows that, on average, applications lose 65% of their users in the first week, while very popular applications (top 100) lose only 35%. It also reveals, however, that many applications have more complex usage behaviour patterns due to seasonality, marketing, or other factors. To capture such effects, we develop a novel app-usage trend measure which provides instantaneous information about the popularity of an application. Analysis of our data using this trend filter shows that roughly 40% of all apps never gain more than a handful of users ( Marginal apps). Less than 0.1% of the remaining 60% are constantly popular ( Dominant apps), 1% have a quick drain of usage after an initial steep rise ( Expired apps), and 6% continuously rise in popularity ( Hot apps). From these, we can distinguish, for instance, trendsetters from copycat apps. We conclude by demonstrating that usage behaviour trend information can be used to develop better mobile app recommendations.
Stephan Sigg, Eemil Lagerspetz, Ella Peltonen, Petteri Nurmi, Sasu Tarkoma
ACM Trans. Web5
2018 GeoMatch: Efficient Large-Scale Map Matching on Apache Spark
abstract
We contribute by developing GeoMatch as a novel, scalable, and efficient big-data pipeline for large-scale map matching on Apache Spark. GeoMatch improves existing spatial big data solutions by utilizing a novel spatial partitioning scheme inspired by Hilbert space-filling curves. Thanks to the partitioning scheme, GeoMatch can effectively balance operations across different processing units and achieve significant performance gains. We demonstrate the effectiveness of GeoMatch through rigorous and extensive benchmarks that consider data sets containing large-scale urban spatial data sets ranging from 166, 253 to 3.78 billion location measurements. Our results show over 17-fold performance improvements compared to previous works while achieving better processing accuracy than current solutions (97.48%).
Ayman Zeidan, Eemil Lagerspetz, Kai Zhao 0011, Petteri Nurmi, Sasu Tarkoma, Huy T. Vo
IEEE BigData5
2018 The hidden image of mobile apps: geographic, demographic, and cultural factors in mobile usage
abstract
While mobile apps have become an integral part of everyday life, little is known about the factors that govern their usage. Particularly the role of geographic and cultural factors has been understudied. This article contributes by carrying out a large-scale analysis of geographic, cultural, and demographic factors in mobile usage. We consider app usage gathered from 25,323 Android users from 44 countries and 54,776 apps in 55 categories, and demographics information collected through a user survey. Our analysis reveals significant differences in app category usage across countries and we show that these differences, to large degree, reflect geographic boundaries. We also demonstrate that country gives more information about application usage than any demographic, but that there also are geographic and socio-economic subgroups in the data. Finally, we demonstrate that app usage correlates with cultural values using the Value Survey Model of Hofstede as a reference of cross-cultural differences.
Ella Peltonen, Eemil Lagerspetz, Jonatan Hamberg, Abhinav Mehrotra, Mirco Musolesi, Petteri Nurmi, Sasu Tarkoma
MobileHCI7
2018 Sensorclone: a framework for harnessing smart devices with virtual sensors
abstract
IoT services hosted by low-power devices rely on the cloud infrastructure to propagate their ubiquitous presence over the Internet. A critical challenge for IoT systems is to ensure continuous provisioning of IoT services by overcoming network breakdowns, hardware failures, and energy constraints. To overcome these issues, we propose a cloud-based framework namely SensorClone, which relies on virtual devices to improve IoT resilience. A virtual device is the digital counterpart of a physical device that has learned to emulate its operations from sample data collected from the physical one. SensorClone exploits the collected data of low-power devices to create virtual devices in the cloud. SensorClone then can opportunistically migrate virtual devices from the cloud into other devices, potentially underutilized, with higher capabilities and closer to the edge of the network, e.g., smart devices. Through a real deployment of our SensorClone in the wild, we identify that virtual devices can be used for two purposes, 1) to reduce the energy consumption of physical devices by duty cycling their service provisioning between the physical device and the virtual representation hosted in the cloud, and 2) to scale IoT services at the edge of the network by harnessing temporal periods of underutilization of smart devices. To evaluate our framework, we present a use case of a virtual sensor created from an IoT service of temperature. From our results, we verify that it is possible to achieve unlimited availability up to 90% and substantial power efficiency under acceptable levels of quality of service. Our work makes contributions towards improving IoT scalability and resilience by using virtual devices.
Huber Flores, Pan Hui 0001, Sasu Tarkoma, Yong Li 0008, Theodoros Anagnostopoulos, Vassilis Kostakos, Chu Luo, Xiang Su 0004
MMSys3
2018 Demo: MegaSense: Megacity-scale Accurate Air Quality Sensing with the Edge
abstract
This demo presents MegaSense, an air pollution monitoring system for realizing low-cost, near real-time and high resolution spatio-temporal air pollution maps of urban areas. MegaSense involves a novel hierarchy of multi-vendor distributed air quality sensors, in which accurate sensors calibrate lower cost sensors. Current low-cost air quality sensors suffer from measurement drift and they have low accuracy. We address this significant open problem for dense urban areas by developing a calibration scheme that detects and automatically corrects drift. MegaSense integrates with the 5G cellular network and leverages mobile edge computing for sensor management and distributed pollution map creation. We demonstrate MegaSense with two sensor types, a state of the art air quality monitoring station and a low-cost sensor array, with calibration between the two to improve the accuracy of the low-cost device. Participants can interact with the sensors and see air quality changes in real-time, and observe the mechanism to mitigate sensor drift. Our re-calibration method minimizes the error for NO2 and O3 81% of the time (vs single calibration) and reduces the mean relative error by 25%-45%.
Eemil Lagerspetz, Samu Varjonen, Francesco Concas, Julien Mineraud, Sasu Tarkoma
MobiCom5
2018 Real-Time IoT Device Activity Detection in Edge Networks
Ibbad Hafeez, Aaron Yi Ding, Markku Antikainen, Sasu Tarkoma
NSS4
2018 Evidence-Aware Mobile Computational Offloading
abstract
Computational offloading can improve user experience of mobile apps through improved responsiveness and reduced energy footprint. A fundamental challenge in offloading is to distinguish situations where offloading is beneficial from those where it is counterproductive. Currently, offloading decisions are predominantly based on profiling performed on individual devices. While significant gains have been shown in benchmarks, these gains rarely translate to real-world use due to the complexity of contexts and parameters that affect offloading. We contribute by proposing crowdsensed evidence traces as a novel mechanism for improving the performance of offloading systems. Instead of limiting to profiling individual devices, crowdsensing enables characterizing execution contexts across a community of users, providing better generalisation and coverage of contexts. We demonstrate the feasibility of using crowdsensing to characterize offloading contexts through an analysis of two crowdsensing datasets. Motivated by our results, we present the design and development of the EMCO toolkit and platform as a novel solution for computational offloading. Experiments carried out on a testbed deployment in Amazon EC2 Ireland demonstrate that EMCO can consistently accelerate app execution while at the same time reduce energy footprint. We also demonstrate that EMCO provides better scalability than current cloud platforms, being able to serve a larger number of clients without variations in performance. Ourframework, use cases, and tools are available as open source from GitHub.
Huber Flores, Pan Hui 0001, Petteri Nurmi, Eemil Lagerspetz, Sasu Tarkoma, Jukka Manner, Vassilis Kostakos, Yong Li 0008, Xiang Su 0004
IEEE Trans. Mob. Comput.5
2018 Secure Cloud Connectivity for Scientific Applications
abstract
Cloud computing improves utilization and flexibility in allocating computing resources while reducing the infrastructural costs. However, in many cases cloud technology is still proprietary and tainted by security issues rooted in the multi-user and hybrid cloud environment. A lack of secure connectivity in a hybrid cloud environment hinders the adaptation of clouds by scientific communities that require scaling-out of the local infrastructure using publicly available resources for large-scale experiments. In this article, we present a case study of the DII-HEP secure cloud infrastructure and propose an approach to securely scale-out a private cloud deployment to public clouds in order to support hybrid cloud scenarios. A challenge in such scenarios is that cloud vendors may offer varying and possibly incompatible ways to isolate and interconnect virtual machines located in different cloud networks. Our approach is tenant driven in the sense that the tenant provides its connectivity mechanism. We provide a qualitative and quantitative analysis of a number of alternatives to solve this problem. We have chosen one of the standardized alternatives, Host Identity Protocol, for further experimentation in a production system because it supports legacy applications in a topologically-independent and secure way.
Lirim Osmani, Salman Zubair Toor, Miika Komu, Matti J. Kortelainen, Tomas Lindén, Rasib Hassan Khan, Paula Eerola, Sasu Tarkoma
IEEE Trans. Serv. Comput.9
2017 A refactoring approach for optimizing mobile networks
abstract
Mobile networks are expected to serve a wide range of verticals, however the Long Term Evolution (LTE) network is optimized for basic mobile operator services only. The network functions serving LTE networks are largely implemented as dedicated single function devices that offer poor customization options. This intrinsic inflexibility makes current LTE networks unable to meet the requirements of future mobile networks. For example, LTE networks experience signaling storms because the signals exchanged by the network functions cannot be optimized according to the current usage pattern of mobile services. Modularizing these network functions would enable a refactoring of the LTE network, allowing operators to compose networks that adapt and evolve with the influx of verticals. In this article, we present a new approach for refactoring the network functions serving LTE networks which can be leveraged to compose a modular mobile network optimized for the verticals using its services. As an example, we demonstrate that deploying network functions at the edge significantly reduces the signals exchanged within a mobile network.
Matteo Pozza, Ashwin Rao, Armir Bujari, Hannu Flinck, Claudio E. Palazzi, Sasu Tarkoma
ICC6
2017 Modeling Mobile Code Acceleration in the Cloud
abstract
The quality of service of a mobile application is critical to ensure user satisfaction. Techniques have been proposed to accomplish adaptation of quality of service dynamically. However, there is still a limited understanding about how to provide a utility model for code execution. One key challenge is modeling the level of quality in the code execution that can be provisioned by the cloud. Since the allocation of cloud resources has a cost, it is important to optimize cloud usage. We propose a software-defined networking approach that allows modeling and controlling code acceleration of a mobile application deployed across multiple type of devices. By segregating the computational requirements of the mobile application into groups, we were able to define the acceleration needed by each group of devices. As the computational requirements of a device can change across time, a mobile device can be re-assigned to another group based on demand. Our SDN approach implements a model that allows the system to predict workload based on acceleration groups. Evaluating our system in a real testbed showed that it is possible to predict workload and allocate optimal resources to handle that workload with 87.5% accuracy.
Huber Flores, Xiang Su 0001, Vassilis Kostakos, Jukka Riekki, Eemil Lagerspetz, Sasu Tarkoma, Pan Hui 0001, Yong Li 0008, Jukka Manner
ICDCS6
2017 IoT SENTINEL: Automated Device-Type Identification for Security Enforcement in IoT
abstract
With the rapid growth of the Internet-of-Things (IoT), concerns about the security of IoT devices have become prominent. Several vendors are producing IP-connected devices for home and small office networks that often suffer from flawed security designs and implementations. They also tend to lack mechanisms for firmware updates or patches that can help eliminate security vulnerabilities. Securing networks where the presence of such vulnerable devices is given, requires a brownfield approach: applying necessary protection measures within the network so that potentially vulnerable devices can coexist without endangering the security of other devices in the same network. In this paper, we present IoT Sentinel, a system capable of automatically identifying the types of devices being connected to an IoT network and enabling enforcement of rules for constraining the communications of vulnerable devices so as to minimize damage resulting from their compromise. We show that IoT Sentinel is effective in identifying device types and has minimal performance overhead.
Markus Miettinen, Samuel Marchal, Ibbad Hafeez, N. Asokan, Ahmad-Reza Sadeghi, Sasu Tarkoma
ICDCS6
2017 IoT Sentinel Demo: Automated Device-Type Identification for Security Enforcement in IoT
abstract
The emergence of numerous new manufacturers producing devices for the Internet-of-Things (IoT) has given rise to new security concerns. Many IoT devices exhibit security flaws making them vulnerable for attacks and manufacturers have difficulties in providing appropriate security patches to their products in a timely and user-friendly manner. In this paper, we present our implementation of IoT Sentinel, which is a system aimed at protecting the user's network from vulnerable IoT devices. IoT Sentinel automatically identifies vulnerable devices when they are first introduced to the network and enforces appropriate traffic filtering rules to protect other devices from the threats originating from the vulnerable devices.
Markus Miettinen, Samuel Marchal, Ibbad Hafeez, Tommaso Frassetto, N. Asokan, Ahmad-Reza Sadeghi, Sasu Tarkoma
ICDCS7
2017 Poster: IoTURVA: Securing Device-to-Device Communications for IoT
abstract
In this poster we present IoTurva, a platform for securing Device-to-Device (D2D) communication in IoT. Our solution takes a blackbox approach to secure IoT edge-networks. We combine user and device-centric context-information together with network data to classify network communication as normal or malicious. We have designed a dual-layer traffic classification scheme based on fuzzy logic, where the classification model is trained remotely. The remotely trained model is then used by the edge gateway to classify the network traffic. We have implemented a proof-of-concept prototype and evaluate its performance in a real world environment. Theevaluation shows that IoTurva causes very small overhead while it works with minimal hardware, and that our model training and classification approach can improve system efficiency and privacy.
Ibbad Hafeez, Aaron Yi Ding, Markku Antikainen, Sasu Tarkoma
MobiCom4
2017 Differentially private Bayesian learning on distributed data
abstract
Many applications of machine learning, for example in health care, would benefit from methods that can guarantee privacy of data subjects. Differential privacy (DP) has become established as a standard for protecting learning results. The standard DP algorithms require a single trusted party to have access to the entire data, which is a clear weakness, or add prohibitive amounts of noise. We consider DP Bayesian learning in a distributed setting, where each party only holds a single sample or a few samples of the data. We propose a learning strategy based on a secure multi-party sum function for aggregating summaries from data holders and the Gaussian mechanism for DP. Our method builds on an asymptotically optimal and practically efficient DP Bayesian inference with rapidly diminishing extra cost.
Mikko A. Heikkilä, Eemil Lagerspetz, Samuel Kaski, Kana Shimizu, Sasu Tarkoma, Antti Honkela
NIPS5
2017 Guest Editorial Privacy Issues in Internet of Things
abstract
The long-heralded Internet of Things (IoT) is finally becoming a reality. From factories and the ubiquitous Internet-connected fridge, we now see heating control systems, cars, dishwashers, and all manner of common-place devices being connected. While this has certainly realized new capabilities, such as the ability to control one’s domestic heating remotely, the benefits are perhaps more mixed: every device that we can remotely control is a device that someone else can remotely hack. And that is just the devices—the literal things; in tandem we also see increasing intrusion of Internet-connectivity into services, practices and everyday infrastructures such as transport and retail. Coupled with the increasingly invasive deployment of these devices into everyday lives, the result is a substantial increase in threats to privacy arising from the IoT.
Richard Mortier, Jon Crowcroft, Charith Perera, Sasu Tarkoma, Peter Christen
IEEE Internet Things J.4
2017 Social-aware hybrid mobile offloading
abstract
Mobile offloading is a promising technique to aid the constrained resources of a mobile device. By offloading a computational task, a device can save energy and increase the performance of the mobile applications. Unfortunately, in existing offloading systems, the opportunistic moments to offload a task are often sporadic and short-lived. We overcome this problem by proposing a social-aware hybrid offloading system (HyMobi), which increases the spectrum of offloading opportunities. As a mobile device is always co-located to at least one source of network infrastructure throughout of the day, by merging cloudlet, device-to-device and remote cloud offloading, we increase the availability of offloading support. Integrating these systems is not trivial. In order to keep such coupling, a strong social catalyst is required to foster user's participation and collaboration. Thus, we equip our system with an incentive mechanism based on credit and reputation, which exploits users’ social aspects to create offload communities. We evaluate our system under controlled and in-the-wild scenarios. With credit, it is possible for a device to create opportunistic moments based on user's present need. As a result, we extended the widely used opportunistic model with a long-term perspective that significantly improves the offloading process and encourages unsupervised offloading adoption in the wild.
Huber Flores, Rajesh Sharma 0002, Denzil Ferreira, Vassilis Kostakos, Jukka Manner, Sasu Tarkoma, Pan Hui 0001, Yong Li 0008
Pervasive Mob. Comput.6
2017 Full Charge Capacity and Charging Diagnosis of Smartphone Batteries
abstract
Full charge capacity (FCC) refers to the amount of charge a battery can hold. It is the fundamental property of smartphone batteries that diminishes as the battery ages and is charged/discharged. We investigate the behavior of smartphone batteries while charging and demonstrate that battery voltage and charging rate information can together characterize the FCC of a battery. We propose a new method for accurately estimating FCC without exposing low-level system details or introducing new hardware or system modules. We further propose and implement a collaborative FCC estimation technique that builds on crowd-sourced battery data. The method finds the reference voltage curve and charging rate of a particular smartphone model from the data and then compares with those of an individual device. After analyzing a large data set towards a crowd-sourced rate versus FCC model, we report that 55 percent of all devices and at least one device in 330 out of 357 unique device models lost some of their FCC. For some old device models, the median capacity loss exceeded 20 percent. The models further enable debugging the performance of smartphone charging. We propose an algorithm, called BatterySense, which utilizes crowd-sourced rate to detect abnormal charging performance, estimate FCC of the device battery, and detect battery changes.
Mohammad Ashraful Hoque, Matti Siekkinen, Jonghoe Koo, Sasu Tarkoma
IEEE Trans. Mob. Comput.4
2016 Too big to mail: On the way to publish large-scale mobile analytics data
abstract
The Carat project started in 2012 has collected over 1.5 TB of data from over 850,000 mobile users all over the world. The project uses Apache Thrift to transmit data, and Apache Spark to run data analysis tasks, and the gist of the Carat analysis method has been published. While the Carat application code is open source, the data is much harder to share because of its size and privacy concerns. This paper outlines the challenges in sharing such a large-scale dataset with detailed information about smart devices, applications, and their users, and presents some solutions to these challenges.
Ella Peltonen, Eemil Lagerspetz, Petteri Nurmi, Sasu Tarkoma
IEEE BigData4
2016 Urban human mobility data mining: An overview
abstract
Understanding urban human mobility is crucial for epidemic control, urban planning, traffic forecasting systems and, more recently, various mobile and network applications. Nowadays, a variety of urban human mobility data have been gathered and published. Pervasive GPS data can be collected by mobile phones. A mobile operator can track people's movement in cities based on their cellular network location. This urban human mobility data contains rich knowledge about locations and can help in addressing many urban challenges such as traffic congestion or air pollution problems. In this article, we survey recent literature on urban human mobility from a data mining view: from the data collection and cleaning, to the mobility models and the applications. First, we summarize recent public urban human mobility data sets and how to clean and preprocess such data. Second, we describe recent urban human mobility models and predictors, e.g., the deep learning predictor, for predicting urban human mobility. Third, we describe how to evaluate the models and predictors. We conclude by considering how applications can utilize the mobility models and predictive tools for addressing city challenges.
Kai Zhao 0011, Sasu Tarkoma, Huy T. Vo
IEEE BigData2
2016 Quantitative evaluation of public spaces using crowd replication
abstract
We propose crowd replication as a low-effort, easy to implement and cost-effective mechanism for quantifying the uses, activities, and sociability of public spaces. Crowd replication combines mobile sensing, direct observation, and mathematical modeling to enable resource efficient and accurate quantification of public spaces. The core idea behind crowd replication is to instrument the researcher investigating a public space with sensors embedded on commodity devices and to engage him/her into imitation of people using the space. By combining the collected sensor data with a direct observations and population model, individual sensor traces can be generalized to capture the behavior of a larger population. We validate the use of crowd replication as a data collection mechanism through a field study conducted within an exemplary metropolitan urban space. Results of our evaluation show that crowd replication accurately captures real human dynamics (0.914 correlation between indicators estimated from crowd replication and visual surveillance) and captures data that is representative of the behavior of people within the public space.
Samuli Hemminki, Keisuke Kuribayashi, Shin'ichi Konomi, Petteri Nurmi, Sasu Tarkoma
SIGSPATIAL/GIS5
2016 Off-the-Shelf Software-defined Wi-Fi Networks
abstract
Wi-Fi networks were one of the first use-cases for Software-defined networking (SDN). However, to deploy a software-defined Wi-Fi network today, one has to rely on research prototypes with availability, documentation, hardware requirements, and scalability issues. To alleviate this situation, we demonstrate two simple techniques to bring SDN functionality to existing Wi-Fi networks and discuss their benefits and short-comings. Researchers can use our techniques to convert their existing Wi-Fi testbeds into software defined Wi-Fi testbeds. Our two techniques thus significantly lower the barrier-to-entry for deploying software-defined Wi-Fi networks.
Seppo Hätönen, Petri Savolainen, Ashwin Rao, Hannu Flinck, Sasu Tarkoma
SIGCOMM5
2016 A gap analysis of Internet-of-Things platforms
Julien Mineraud, Oleksiy Mazhelis, Xiang Su 0001, Sasu Tarkoma
Comput. Commun.4
2016 Constella: Crowdsourced system setting recommendations for mobile devices
Ella Peltonen, Eemil Lagerspetz, Petteri Nurmi, Sasu Tarkoma
Pervasive Mob. Comput.4
2015 Checksum gestures: continuous gestures as an out-of-band channel for secure pairing
abstract
We propose the use of a single continuous gesture as a novel, intuitive, and efficient mechanism to authenticate a secure communication channel. Our approach builds on a novel algorithm for encoding (at least 20-bits) authentication information as a single continuous gesture, referred to as a checksum gesture. By asking the user to perform the generated gesture, a secure channel can be authenticated. Results from a controlled user experiment (N = 13 participants, 1022 trials) demonstrate the feasibility of our technique, showing over 90% success rate in establishing a secure communication channel despite relying on complex gesture patterns. The authentication times of our method are over three-folds faster than with previous gesture-based solutions. The average execution time of a gesture is 5:7 seconds in our study, which is comparable to the input time of conventional text input based PIN authentication. Our approach is particularly well-suited for scenarios involving wearable devices that lack conventional input capabilities, e.g., pairing a smartwatch with an interactive display.
Imtiaj Ahmed, Yina Ye, Sourav Bhattacharya, N. Asokan, Giulio Jacucci, Petteri Nurmi, Sasu Tarkoma
UbiComp7
2015 Towards disruption tolerant ICN
abstract
Information-Centric Networking (ICN) is a prominent topic in current networking research. ICN design significantly considers the increased demand of scalable and efficient content distribution for Future Internet. However, intermittently connected mobile environments or disruptive networks present a significant challenge to ICN deployment. In this context, delay tolerant networking (DTN) architecture is an initiative that effectively deals with network disruptions. Among all ICN proposals, Content Centric Networking (CCN) is gaining more and more interest for its architectural design, but still has the limitation in highly disruptive environment. In this paper, we design a protocol stack referred as CCNDTN which integrates DTN architecture in the native CCN to deal with network disruption. We also present the implementation details of the proposed CCNDTN. We extend CCN routing strategies by integrating Bundle protocol of DTN architecture. The integration of CCN and DTN enriches the connectivity options of CCN architecture in fragmented networks. Furthermore, CCNDTN can be beneficial through the simultaneous use of all available connectivities and opportunistic networking of DTN for the dissemination of larger data items. This paper also highlights the potential use cases of CCNDTN architecture and crucial questions about integrating CCN and DTN.
Hasan M. A. Islam, Andrey Lukyanenko, Sasu Tarkoma, Antti Ylä-Jääski
ISCC3
2015 Demo: An Open-source Software Defined Platform for Collaborative and Energy-aware WiFi Offloading
abstract
This demonstration presents a novel software defined platform for achieving collaborative and energy-aware WiFi offloading. The platform consists of an extensible central controller, programmable offloading agents, and offloading extensions on mobile devices. Driven by our extensive measurements of energy consumption on smartphones, we propose an effective energy-aware offloading algorithm and integrate it to our platform. By enabling collaboration between wireless networks and mobile users, our solution can make optimal offloading decisions that improve offloading efficiency for network operators and achieve energy saving for mobile users. To enhance deployability, we have released our platform under open-source licenses on GitHub.
Aaron Yi Ding, Yanhe Liu, Sasu Tarkoma, Hannu Flinck, Jon Crowcroft
MobiCom3
2015 Poster: VPN Tunnels for Energy Efficient Multimedia Streaming
abstract
Minimizing the energy consumption of mobile devices for wireless network access is important. In this article, we analyze the energy efficiency of a new set of applications which use Virtual Private Network (VPN) tunnels for secure communication. First, we discuss the energy efficiency of a number of VPN applications from a large scale deployment of 500 K devices. We next measure the energy consumption of some of these applications with different use cases. Finally, we demonstrate that a VPN tunnel can be instrumented for enhanced energy efficiency with multimedia streaming applications. Our results indicate energy savings of 40% for this class of applications.
Mohammad Ashraful Hoque, Kasperi Saarikoski, Eemil Lagerspetz, Julien Mineraud, Sasu Tarkoma
MobiCom5
2015 Poster: Extremely Parallel Resource Pre-Fetching for Energy Optimized Mobile Web Browsing
abstract
Mobile web browsing is experienced slow because of the limited rendering capability of the mobile devices, wireless latency, and incremental rendering of the page or resource loading. The browser renders resources in between two consecutive resource downloads. However, during this period, the wireless interfaces consume energy doing nothing useful. In this work, we measure the performance of SPDY for mobile web browsing. We demonstrate that mobile devices waste energy by keeping the wireless network interface idle between consecutive resource downloads. We next show that by identifying the embedded resources in a web page and downloading those resources in parallel at the very beginning can reduce the small idle periods and thus energy consumption by 20-50%, depending on the wireless network type.
Mohammad Ashraful Hoque, Sasu Tarkoma, Tuikku Anttila
MobiCom2
2015 Energy modeling of system settings: A crowdsourced approach
abstract
The question “Where has my battery life gone?” remains a common source of frustration for many smartphone users. With the increased complexity of smartphone applications, and the increasing number of system settings affecting them, understanding and optimizing battery use has become a difficult chore. The present paper develops a novel approach for constructing energy models from crowdsourced measurements. In contrast to previous approaches, which have focused on the effect of a specific sensor, system setting or application, our approach can simultaneously capture relationships between multiple factors, and provide a unified view of the energy state of the mobile device. We demonstrate the validity of using crowdsourced measurements for constructing battery models through a combination of large-scale analysis of a dataset containing battery discharge and system state measurements and hardware power measurements. The results indicate that the models captured by our approach are both in line with previous studies on battery consumption and empirical measurements, providing a cost-effective way to construct energy models during normal operations of the device. The analysis also provides several new insights about battery consumption. For example, our analysis shows the energy use of high CPU activity with automatic screen brightness is actually higher (resulting in around 9 minutes shorter battery lifetime on average) than with a medium CPU load and manual screen brightness; a Wi-Fi signal strength drop of one bar can result in a battery life loss of over 13%; and a smartphone sitting in the sun can experience over 50% worse battery life than one indoors in cool conditions.
Ella Peltonen, Eemil Lagerspetz, Petteri Nurmi, Sasu Tarkoma
PerCom4
2015 Blending Problem- and Project-Based Learning in Internet of Things Education: Case Greenhouse Maintenance
abstract
This article describes an experimental course where students developed Internet of Things device prototypes to improve the upkeep of an urban rooftop greenhouse. With the help of a problem-based learning approach, students were first familiarized with their new learning environment and encouraged to find issues that could be improved as a meaningful personal learning experience. A project-based learning approach was then used to develop innovative solutions while validating their relevance in collaboration with gardeners that were taking care of the greenhouse. As a result, a number of practical applications for monitoring the state of the greenhouse were developed along with new practices for its maintenance. As participants were given the freedom to choose both the topic and technologies to work with, the course provided a learning experience that was tailored to suit personal interests and competences. Having the common background story allowed students to practice teamwork skills and collaborative software engineering in the context of the emerging topic of Internet of Things.
Hanna Mäenpää, Sasu Tarkoma, Samu Varjonen, Arto Vihavainen
SIGCSE2
2015 Mobile multimedia streaming techniques: QoE and energy saving perspective
Mohammad Ashraful Hoque, Matti Siekkinen, Jukka K. Nurminen, Mika Aalto, Sasu Tarkoma
Pervasive Mob. Comput.5
2015 Towards Maximizing Timely Content Delivery in Delay Tolerant Networks
abstract
Many applications, such as product promotion advertisement and traffic congestion notification, benefit from opportunistic content exchange in Delay Tolerant Networks (DTNs). An important requirement of such applications is timely delivery. However, the intermittent connectivity of DTNs may significantly delay content exchange, and cannot guarantee timely delivery. The state-of-the-arts capture mobility patterns or social properties of mobile devices. Such solutions do not capture patterns of delivered content in order to optimize content delivery. Without such optimization, the content demanded by a large number of subscribers could follow the same forwarding path as the content by only one subscriber, leading to traffic congestion and packet drop. To address the challenge, in this paper, we develop a solution framework, namely Ameba, for timely delivery. In detail, we first leverage content properties to derive an optimal routing hop count of each content to maximize the number of needed nodes. Next, we develop node utilities to capture interests, capacity and locations of mobile devices. Finally, the distributed forwarding scheme leverages the optimal routing hop count and node utilities to deliver content towards the needed nodes in a timely manner. Illustrative results verify that Ameba achieves comparable delivery ratio as Epidemic but with much lower overhead.
Weixiong Rao, Kai Zhao 0011, Yan Zhang 0002, Pan Hui 0001, Sasu Tarkoma
IEEE Trans. Mob. Comput.5
2015 MTAF: An Adaptive Design for Keyword-Based Content Dissemination on DHT Networks
abstract
Beyond offering the widely used keyword search function, many peer-to-peer systems nowadays support the subscription function. For example, Vuze allows users to create subscription filters based on the keyword search. Given the subscription, episodic or related content will be delivered to the users whenever new episodes are available. Unfortunately, these applications suffer from the downsides, for example, high network traffic in the nodes maintaining popular terms. In this paper, we propose the MTAF mechanism to overcome the issues. The key of MTAF is to carefully select a subset of terms without incurring false negatives and to forward the content item toward the home nodes of such selected terms for low content forwarding cost. Experimental results based on real datasets indicate that the proposed solutions are efficient compared to existing approaches. In particular, the similarity-based replication of filters is shown to mitigate the effect of hot spots that arise due to the fact that some document terms are substantially more popular than the others.
Weixiong Rao, Roman Vitenberg, Lei Chen 0002, Sasu Tarkoma
IEEE Trans. Parallel Distributed Syst.4
2014 How carat affects user behavior: implications for mobile battery awareness applications
abstract
Mobile devices have limited battery life, and numerous battery management applications are available that aim to improve it. This paper examines a large-scale mobile battery awareness application, called Carat, to see how it changes user behavior with long-term use. We conducted a survey of current Carat Android users and analyzed their interaction logs. The results show that long-term Carat users save more battery, charge their devices less often, learn to manage their battery with less help from Carat, have a better understanding of how Carat works, and may enjoy competing against other users. Based on these findings, we propose a set of guidelines for mobile battery awareness applications: battery awareness applications should make the reasoning behind their recommendations understandable to the user, be tailored to retain long-term users, take the audience into account when formulating feedback, and distinguish third-party and system applications.
Kumaripaba Athukorala, Eemil Lagerspetz, Maria von Kügelgen, Antti Jylhä, Adam J. Oliner, Sasu Tarkoma, Giulio Jacucci
CHI6
2014 Gravity and linear acceleration estimation on mobile devices
abstract
Linear acceleration is an important enabler for many applications of mobile and wearable activity recognition. The most common approach for estimating linear acceleration is to estimate the gravity component of accelerometer measurements and to project gravity-eliminated accelerometer measurements o
Samuli Hemminki, Petteri Nurmi, Sasu Tarkoma
MobiQuitous3
2014 Poster: SoftOffload: a programmable approach toward collaborative mobile traffic offloading
abstract
The fast increase of mobile traffic from smartphone-like devices has created a huge pressure for the cellular operators to manage the network infrastructure and resources. Offloading the mobile traffic to alternative networks such as WiFi is sought as a promising direction to solve this problem cost-effectively. According to our study and experimental findings, existing research proposals are lack of concern for the complexity of network deployment and device limitations, which impedes the solution deployment. To overcome such challenge, we propose SoftOffload, a programmable framework for collaborative mobile traffic offloading. SoftOffload takes the advantage of software defined networking (SDN) paradigm in terms of openness and extensibility. We have implemented the first prototype utilising the open source Floodlight platform.
Aaron Yi Ding, Jon Crowcroft, Sasu Tarkoma
MobiSys3
2014 The company you keep: mobile malware infection rates and inexpensive risk indicators
abstract
There is little information from independent sources in the public domain about mobile malware infection rates. The only previous independent estimate (0.0009%) [11], was based on indirect measurements obtained from domain-name resolution traces. In this paper, we present the first independent study of malware infection rates and associated risk factors using data collected directly from over 55,000 Android devices. We find that the malware infection rates in Android devices estimated using two malware datasets (0.28% and 0.26%), though small, are significantly higher than the previous independent estimate. Based on the hypothesis that some application stores have a greater density of malicious applications and that advertising within applications and cross-promotional deals may act as infection vectors, we investigate whether the set of applications used on a device can serve as an indicator for infection of that device. Our analysis indicates that, while not an accurate indicator of infection by itself, the application set does serve as an inexpensive method for identifying the pool of devices on which more expensive monitoring and analysis mechanisms should be deployed. Using our two malware datasets we show that this indicator performs up to about five times better at identifying infected devices than the baseline of random checks. Such indicators can be used, for example, in the search for new or previously undetected malware. It is therefore a technique that can complement standard malware scanning. Our analysis also demonstrates a marginally significant difference in battery use between infected and clean devices.
Hien Thi Thu Truong, Eemil Lagerspetz, Petteri Nurmi, Adam J. Oliner, Sasu Tarkoma, N. Asokan, Sourav Bhattacharya
WWW5
2014 Software defined networking for security enhancement in wireless mobile networks
Aaron Yi Ding, Jon Crowcroft, Sasu Tarkoma, Hannu Flinck
Comput. Networks3
2014 Tolerating path heterogeneity in multipath TCP with bounded receive buffers
Ming Li 0035, Andrey Lukyanenko, Sasu Tarkoma, Yong Cui 0001, Antti Ylä-Jääski
Comput. Networks3
2014 Modeling Energy Consumption of Data Transmission Over Wi-Fi
abstract
Wireless data transmission consumes a significant part of the overall energy consumption of smartphones, due to the popularity of Internet applications. In this paper, we investigate the energy consumption characteristics of data transmission over Wi-Fi, focusing on the effect of Internet flow characteristics and network environment. We present deterministic models that describe the energy consumption of Wi-Fi data transmission with traffic burstiness, network performance metrics like throughput and retransmission rate, and parameters of the power saving mechanisms in use. Our models are practical because their inputs are easily available on mobile platforms without modifying low-level software or hardware components. We demonstrate the practice of model-based energy profiling on Maemo, Symbian, and Android phones, and evaluate the accuracy with physical power measurement of applications including file transfer, web browsing, video streaming, and instant messaging. Our experimental results show that our models are of adequate accuracy for energy profiling and are easy to apply.
Yu Xiao 0001, Yong Cui 0001, Petri Savolainen, Matti Siekkinen, Antti Ylä-Jääski, Sasu Tarkoma
IEEE Trans. Mob. Comput.8
2014 Saving Energy in Mobile Devices for On-Demand Multimedia Streaming - A Cross-Layer Approach
abstract
This article proposes a novel energy-efficient multimedia delivery system called EStreamer. First, we study the relationship between buffer size at the client, burst-shaped TCP-based multimedia traffic, and energy consumption of wireless network interfaces in smartphones. Based on the study, we design and implement EStreamer for constant bit rate and rate-adaptive streaming. EStreamer can improve battery lifetime by 3x, 1.5x, and 2x while streaming over Wi-Fi, 3G, and 4G, respectively.
Mohammad Ashraful Hoque, Matti Siekkinen, Jukka K. Nurminen, Sasu Tarkoma, Mika Aalto
ACM Trans. Multim. Comput. Commun. Appl.4
2014 Evaluating continuous top-k queries over document streams
Weixiong Rao, Lei Chen 0002, Shudong Chen, Sasu Tarkoma
World Wide Web4
2013 Subscription Privacy Protection in Topic-Based Pub/Sub
Weixiong Rao, Lei Chen 0002, Mingxuan Yuan, Sasu Tarkoma, Hong Mei 0001
DASFAA (1)4
2013 The Delayed ACK evolution in MPTCP
abstract
Multipath TCP (MPTCP) is a major extension of TCP that aims to offer higher aggregate bandwidth and robustness by pooling multiple paths within one transport connection. One of the current weaknesses of the protocol is the Delayed ACK scheme inherited from TCP. The Delayed ACK is an option of TCP that allows the receiver to delay sending an ACK for every other packet within a window given by the Delayed ACK timer. At the sender, the RTO should be no less than the Minimum RTO to avoid spurious timeouts. This strategy can lead to significant performance degradation in the presence of timeouts, especially in high speed networks, where RTT is usually one- or two-order of magnitude smaller than the Minimum RTO. When a subflow occurs a timeout, the receiver has to buffer data from all the subflows until the missing packet is received. The data may overrun the receive buffer to cause flow control at the sender, which seriously impacts the overall performance. In order to avoid MPTCP performance degradation, we propose a new Delayed ACK aiming to remove the Minimum RTO constraint at the sender while to reserve the Delayed ACK function at the receiver. Our solution requires only minor modification to the legacy Delayed ACK scheme and it introduces negligible computational overhead and no extra traffic overhead. We use a NS-3 network simulator to evaluate the performance of our new Delayed ACK in three typical network environments. The results indicate that MPTCP using new Delayed ACK scheme requires much smaller aggregate buffer than MPTCP using the legacy Delayed ACK scheme, especially in high speed networks where the buffer requirement reduces one- or two-order of magnitude.
Ming Li 0035, Andrey Lukyanenko, Sasu Tarkoma, Antti Ylä-Jääski
GLOBECOM3
2013 Replication-Based Load Balancing in Distributed Content-Based Publish/Subscribe
abstract
In recent years, content-based publish/subscribe (pub/sub) has become a popular paradigm to decouple content producers and consumers for Internet-scale content services. Many real applications show that the content workloads frequently exhibit very skewed distribution, and incur unbalanced workloads. To balance the workloads, the literature of content-based pub/sub adopted a migration scheme (Mis) to move (a subset of) subscription filters from overloaded brokers to underloaded brokers. In this way, the publications that successfully match the moved filters are then overloaded, leading to balanced workloads. Unfortunately, the scheme cannot reduce the overall matching workloads. In the worst case, suppose that all brokers suffer from heavy workloads. cannot find available brokers to offload the heavy workloads of those overloaded brokers, and fails to balance the workloads. To overcome the issue, the contribution of this paper is to develop a set of novel load balancing algorithms, namely a similarity-based replication scheme (Sir). The novelty of is that it not only balances the workloads of brokers but also reduces the overall workloads. Based on both simulation and emulation results, the extensive experiments verify that can achieve much better performance than, in terms of 43.10% higher entropy value (i.e., more balanced workloads) and 46.39% lower workloads.
Weixiong Rao, Pan Hui 0001, Sasu Tarkoma
IPDPS4
2013 Spaceify: a client-edge-server ecosystem for mobile computing in smart spaces
abstract
Spaceify is a novel edge architecture and an ecosystem for smart spaces --- a technology that extends the mobile user view of today's common space services (e.g., WiFi) to a richer portfolio of space-centric, localized services and space-interactive applications.
Petri Savolainen, Abdelsalam Helal, Jukka Reitmaa, Kai Kuikkaniemi, Giulio Jacucci, Mikko Rinne, Marko Turpeinen, Sasu Tarkoma
MobiCom8
2013 Efficient new delayed ACK for TCP: old problem, new insight
abstract
When a TCP connection experiences a timeout, the sender must wait at least RTOmin (Minimum Retransmission Timeout) before doing the retransmission, during which the channel may be completely idle, undermining the throughput and channel efficiency. In this paper, we investigate the origin of RTOmin and find that it is needed to mitigate against spurious timeouts when the Delayed ACK (DA) scheme for TCP is implemented.
Ming Li 0035, Andrey Lukyanenko, Sasu Tarkoma, Antti Ylä-Jääski
MSWiM3
2013 Enabling energy-aware collaborative mobile data offloading for smartphones
abstract
Searching for mobile data offloading solutions has been topical in recent years. In this paper, we present a collaborative WiFi-based mobile data offloading architecture - Metropolitan Advanced Delivery Network (MADNet), targeting at improving the energy efficiency for smartphones. According to our measurements,WiFi-based mobile data offloading for moving smartphones is challenging due to the limitation ofWiFi antennas deployed on existing smartphones and the short contact duration with WiFi APs. Moreover, our study shows that the number of open-accessible WiFi APs is very limited for smartphones in metropolitan areas, which significantly affects the offloading opportunities for previous schemes that use only open APs. To address these problems, MADNet intelligently aggregates the collaborative power of cellular operators, WiFi service providers and end-users. We design an energy-aware algorithm for energy-constrained devices to assist the offloading decision. Our design enables smartphones to select the most energy efficient WiFi AP for offloading. The experimental evaluation of our prototype on smartphone (Nokia N900) demonstrates that we are able to achieve more than 80% energy saving. Our measurement results also show that MADNet can tolerate minor errors in localization, mobility prediction, and offloading capacity estimation.
Aaron Yi Ding, Bo Han 0001, Yu Xiao 0001, Pan Hui 0001, Aravind Srinivasan, Markku Kojo, Sasu Tarkoma
SECON7
2013 Accelerometer-based transportation mode detection on smartphones
abstract
We present novel accelerometer-based techniques for accurate and fine-grained detection of transportation modes on smartphones. The primary contributions of our work are an improved algorithm for estimating the gravity component of accelerometer measurements, a novel set of accelerometer features that are able to capture key characteristics of vehicular movement patterns, and a hierarchical decomposition of the detection task. We evaluate our approach using over 150 hours of transportation data, which has been collected from 4 different countries and 16 individuals. Results of the evaluation demonstrate that our approach is able to improve transportation mode detection by over 20% compared to current accelerometer-based systems, while at the same time improving generalization and robustness of the detection. The main performance improvements are obtained for motorised transportation modalities, which currently represent the main challenge for smartphone-based transportation mode detection.
Samuli Hemminki, Petteri Nurmi, Sasu Tarkoma
SenSys3
2013 CoSense: a collaborative sensing platform for mobile devices
abstract
We introduce CoSense, a collaborative sensing platform for mobile devices that opportunistically distributes sensing tasks between familiar devices in close proximity. We use empirical energy measurements together with data collected from everyday transportation behaviour to demonstrate that our solution can significantly reduce power consumption while maintaining the best possible sensing accuracy.
Samuli Hemminki, Kai Zhao 0011, Aaron Yi Ding, Martti Rannanjärvi, Sasu Tarkoma, Petteri Nurmi
SenSys5
2013 Carat: collaborative energy diagnosis for mobile devices
abstract
We aim to detect and diagnose energy anomalies, abnormally heavy battery use. This paper describes a collaborative black-box method, and an implementation called Carat, for diagnosing anomalies on mobile devices. A client app sends intermittent, coarse-grained measurements to a server, which correlates higher expected energy use with client properties like the running apps, device model, and operating system. The analysis quantifies the error and confidence associated with a diagnosis, suggests actions the user could take to improve battery life, and projects the amount of improvement. During a deployment to a community of more than 500,000 devices, Carat diagnosed thousands of energy anomalies in the wild. Carat detected all synthetically injected anomalies, produced no known instances of false positives, projected the battery impact of anomalies with 95% accuracy, and, on average, increased a user's battery life by 11% after 10 days (compared with 1.9% for the control group).
Adam J. Oliner, Anand Padmanabha Iyer, Ion Stoica, Eemil Lagerspetz, Sasu Tarkoma
SenSys5
2013 Tolerating path heterogeneity in multipath TCP with bounded receive buffers
abstract
No abstract available.
Ming Li 0035, Andrey Lukyanenko, Sasu Tarkoma, Yong Cui 0001, Antti Ylä-Jääski
SIGMETRICS3
2013 Bitlist: New Full-text Index for Low Space Cost and Efficient Keyword Search
abstract
Nowadays Web search engines are experiencing significant performance challenges caused by a huge amount of Web pages and increasingly larger number of Web users. The key issue for addressing these challenges is to design a compact structure which can index Web documents with low space and meanwhile process keyword search very fast. Unfortunately, the current solutions typically separate the space optimization from the search improvement. As a result, such solutions either save space yet with search inefficiency, or allow fast keyword search but with huge space requirement. In this paper, to address the challenges, we propose a novel structure bitlist with both low space requirement and supporting fast keyword search. Specifically, based on a simple and yet very efficient encoding scheme, bitlist uses a single number to encode a set of integer document IDs for low space, and adopts fast bitwise operations for very efficient boolean-based keyword search. Our extensive experimental results on real and synthetic data sets verify that bitlist outperforms the recent proposed solution, inverted list compression [23, 22] by spending 36.71% less space and 61.91% faster processing time, and achieves comparable running time as [8] but with significantly lower space.
Weixiong Rao, Lei Chen 0002, Pan Hui 0001, Sasu Tarkoma
Proc. VLDB Endow.4
2013 Toward Efficient Filter Privacy-Aware Content-Based Pub/Sub Systems
abstract
In recent years, the content-based publish/subscribe [12], [22] has become a popular paradigm to decouple information producers and consumers with the help of brokers. Unfortunately, when users register their personal interests to the brokers, the privacy pertaining to filters defined by honest subscribers could be easily exposed by untrusted brokers, and this situation is further aggravated by the collusion attack between untrusted brokers and compromised subscribers. To protect the filter privacy, we introduce an anonymizer engine to separate the roles of brokers into two parts, and adapt the k-anonymity and `-diversity models to the contentbased pub/sub. When the anonymization model is applied to protect the filter privacy, there is an inherent tradeoff between the anonymization level and the publication redundancy. By leveraging partial-order-based generalization of filters to track filters satisfying k-anonymity and ℓ-diversity, we design algorithms to minimize the publication redundancy. Our experiments show the proposed scheme, when compared with studied counterparts, has smaller forwarding cost while achieving comparable attack resilience.
Weixiong Rao, Lei Chen 0002, Sasu Tarkoma
IEEE Trans. Knowl. Data Eng.3
2012 MOVE: A Large Scale Keyword-Based Content Filtering and Dissemination System
abstract
The Web 2.0 era is characterized by the emergence of a very large amount of live content. A real time and fine grained content filtering approach can precisely keep users up-to-date the information that they are interested. The key of the approach is to offer a scalable match algorithm. One might treat the content match as a special kind of content search, and resort to the classic algorithm [5]. However, due to blind flooding, [5] cannot be simply adapted for scalable content match. To increase the throughput of scalable match, we propose an adaptive approach to allocate (i.e, replicate and partition) filters. The allocation is based on our observation on real datasets: most users prefer to use short queries, consisting of around 2-3 terms per query, and web content typically contains tens and even thousands of terms per article. Thus, by reducing the number of processed documents, we can reduce the latency of matching large articles with filters, and have chance to achieve higher throughput. We implement our approach on an open source project, Apache Cassandra. The experiment with real datasets shows that our approach can achieve around folds of better throughput than two counterpart state-of-the-arts solutions.
Weixiong Rao, Lei Chen 0002, Pan Hui 0001, Sasu Tarkoma
ICDCS4
2012 Maximizing timely content advertising in DTNs
abstract
Many applications, such as product promotion advertisement and traffic congestion notification, benefit from the opportunistic content exchange in Delay Tolerant Networks (DTNs). An important requirement of such applications is timely delivery. However, the intermittent connectivity of DTNs may significantly delay content exchange and cannot guarantee timely delivery. The state-of-the-arts capture the mobility patterns or social properties of mobile devices. However, there is little optimization in terms of the delivered content. Without such optimization, the content demanded by a large number of subscribers could follow the same forwarding path as the content by only one subscriber. To address the challenge, in this paper, we separate content routing from content forwarding. For content routing, we leverage content properties to derive an optimal routing hop count for each content in order to maximize the number of nodes which receive demanded content. Next, for timely forwarding, we develop node utilities to capture interests and mobility patterns of mobile devices for the selection of content carriers. The distributed greedy relay scheme, Ameba, leverages the optimal routing hop count and developed utilities to timely relay content to the needed nodes as fast as possible. Illustrative results show that Ameba is able to achieve comparable delivery ratio as the Epidemic but with much lower overhead.
Weixiong Rao, Kai Zhao 0011, Yan Zhans, Pan Hui 0001, Sasu Tarkoma
SECON5
2011 Leasing Service for Networks of Interactive Public Displays in Urban Spaces
Marko Jurmu, Hannu Kukka, Simo Hosio, Jukka Riekki, Sasu Tarkoma
GPC5
2011 Towards optimal keyword-based content dissemination in DHT-based P2P networks
abstract
Keyword-based content alert services, e.g., Google Alerts and Microsoft Live Alerts, empower the end users with the ability to automatically receive useful and most recent content. In this paper, we leverage the favorable properties of DHTs, such as scalability, and propose a design of a scalable keyword-based content alert service. The DHT-based architecture matches textual documents with queries based on document terms: For each term, the implementation assigns a home node that is responsible for handling documents and queries that contain the term. The main challenge of this keyword-based matching scheme is the high number of terms that appear in a typical document resulting in a high publication cost. Fortunately, a document can be forwarded to the home nodes of a carefully selected subset of terms without incurring false negatives. In this paper we focus on the MTAF problem of minimizing the number of selected terms to forward the published content. We show that the problem is NP-hardness, and consider centralized and DHT-based solutions. Experimental results based on real datasets indicate that the proposed solutions are efficient compared to existing approaches. In particular, the similarity-based replication of filters that is a key element of our solution is shown to mitigate the effect of hotspots that arise due to the fact that some document terms are substantially more popular than the others, both inside documents and queries.
Weixiong Rao, Roman Vitenberg, Sasu Tarkoma
Peer-to-Peer Computing3
2010 P2P Video-on-Demand: Steady State and Scalability
abstract
The fundamental P2P principle that downloading peers help other peers can be applied in the context of video-on-demand. This represents a demanding application combining aspects of other well-known P2P applications, i.e., live streaming and traditional file sharing. We seek to provide insight on fundamental questions about the performance and scalability of the system. A deterministic fluid model is derived that explicitly takes into account the video transfer and playback phases. The analytical results are complemented with extensive simulations from the corresponding stochastic model, as well as traces from a more realistic BitTorrent simulator.
Samuli Aalto, Pasi E. Lassila, Niklas Raatikainen, Petri Savolainen, Sasu Tarkoma
GLOBECOM5
2010 Cryptographic signatures on the network layer - an alternative to the ISP data retention
abstract
Insecurity of the Internet has led to data retention legislations where user's private data is stored for months or years. Such an approach has significant cost, privacy and security issues. In this paper we propose an alternative way for providing the security and accountability on the Internet by using the Packet Level Authentication (PLA) protocol and perpacket cryptographic signatures. We examine security and privacy properties of our solution. Our analysis shows that using cryptographic identities and signatures on the network level removes the need for costly data retention and actually improves the privacy of users.
Dmitrij Lagutin, Sasu Tarkoma
ISCC2
2010 Segment Level Authentication: Combating internet source spoofing
abstract
This paper presents SLA (Segment Level Authentication), a transport segment level solution designed to prevent both of the intra-domain and inter-domain source spoofing. SLA is based on public key cryptography authentication. It enables intermediate network nodes the ability to validate the packet authenticity by verifying authentication information carried in packets. Although public key cryptography is computationally intensive and induces the traffic overhead, SLA leverages FPGA (Field Programmable Gate Array) based ECC (Elliptic Curve Cryptography) hardware cryptography accelerator to decrease the computation and traffic overhead. SLA provides incremental deployment and offers incentives for both of hosts and ASes. We find that the SLA is feasible for Gigabit links and can effectively mitigate source spoofing in both of intra-domain and inter-domain networks.
Ming Li 0035, Matti Siekkinen, Sasu Tarkoma, Antti Ylä-Jääski, Yong Cui 0001
ISCC3
2010 Dessy: Search and Synchronization on the Move
abstract
Current smartphones have a storage capacity of several gigabytes. More and more information is stored on mobile devices. To meet the challenge of information organization, we turn to desktop search. Users often possess multiple devices, and synchronize (subsets of) information between them. This makes file synchronization more important. This paper presents Dessy, a desktop search and synchronization framework for mobile devices. Dessy supports synchronization of search results, individual files, and directory trees. It allows finding and synchronizing files that reside on remote computers, or the Internet. The contributions of this paper include an energy usage evaluation of the system. Dessy is closely integrated with the Syxaw file synchronizer, which provides efficient file and metadata synchronization, optimizing network usage.
Eemil Lagerspetz, Sasu Tarkoma, Tancred Lindholm
Mobile Data Management2
2010 Dessy: Demonstrating Mobile Search and Synchronization
abstract
The storage capacity of smartphones has reached tens of gigabytes, while the search functionality remains simple. We have designed a search and synchronization framework for mobile devices, called Dessy. Dessy has been designed with mobility and device constraints in mind. It requires only MIDP 2.0 Mobile Java with File Connection support, and Java 1.5 on desktop machines. This paper demonstrates the application in practice, using multiple devices and synchronizing files between a desktop computer, a laptop, smartphones, and the Internet. Smartphones and laptops are able to search for files hosted on other devices as well as on the Internet. Unnecessary search operations are avoided using Bloom filters.
Eemil Lagerspetz, Sasu Tarkoma, Tancred Lindholm
Mobile Data Management2
2009 Flexible Single Sign-On for SIP: Bridging the Identity Chasm
abstract
Identity federation is a key requirement for today's distributed services. This technology allows managed sharing of users' identity information between identity providers (IDP), and subsequently, the use of federated identities to access service providers (SP). Single sign-on (SSO) is a core feature provided by these systems. The Session Initiation Protocol (SIP) is a signaling framework for session call control. It is becoming a widely accepted layer for applications and services, especially in the telecommunications and multimedia domain. In this paper, we explore solutions to incorporate SSO process into the SIP framework in order to simplify the services and resources access. Our design leverages the liberty alliance specifications and extends the existing SIP standards to support SSO functionality. We also present a prototype implementation at the end of this paper.
Pin Nie, Juha-Matti Tapio, Sasu Tarkoma, Jani Heikkinen
ICC3
2009 An Approach to Achieve Context-aware Maps: Combining Semantic Web Technology with Sensor Data
abstract
In this paper, we present our work towards achieving context-awareness in mobile devices by combining Semantic Web technology with sensory data. Our investigation shows that some context data pertaining to the user, such as location, time, and physical surroundings, is vital for the realization of intelligent maps. Hence, embedding context-awareness into intelligent maps may prompt the usability of mobile map applications. Aiming at this goal, we suggest Semantic Web technology-based solution. We present a data representation, Entity Notation, to connect sensors to Resource Description Framework (RDF), the basis of Semantic Web data, in the data interchange level. At the same time, our data representation is lightweight enough that any resource-constrained sensors can support and process it. Ontology and ontology-based inference engine are developed to reason on the sensory data. Finally, intelligent maps could utilize its inference output to achieve context-awareness. We demonstrate our methods with a simulator and discuss the future work.
Xiang Su 0001, Jukka Riekki, Sasu Tarkoma
Intelligent Environments3
2009 Probabilistic routing for multiple flows in wireless multi-hop networks
abstract
Maximizing network throughput is one of the main objectives in wireless multi-hop networks. However, small link bandwidth and severe interference become two great obstacles in network throughput improvement. Single route always encounters great congestion, especially for multiple flows with the same source and destination. Multi-path routing can solve the bandwidth shortage issue to some extent, but conventional routing algorithms of this type have some limitations during route selection and using, which makes bandwidth improvement and interference reduction unfeasible at the same time. In this paper we propose a new metric called effective bandwidth to describe the real bandwidth that a flow can get through a specific route. Based on this metric, we present a probabilistic multi-path routing protocol for multiple flows, which can improve the network throughput by selecting route with bigger effective bandwidth using higher probability. The simulation results show that probabilistic multi-path routing has high superiority over single-path routing in improving network throughput.
Yong Cui 0001, Sasu Tarkoma, Antti Ylä-Jääski
LCN3
2009 Syxaw: Data Synchronization Middleware for the Mobile Web
Tancred Lindholm, Jaakko Kangasharju, Sasu Tarkoma
Mob. Networks Appl.3
2008 Incentive-compatible caching and peering in data-oriented networks
abstract
Several new, data-oriented internetworking architectures have been proposed recently. However, the practical deployability of such designs is an open question. In this paper, we consider data-oriented network designs in the light of the policy and incentive structures of the present internetworking economy. A main observation of our work is that none of the present proposals is both policy-compliant and incentive-compatible with the current internetworking market, which makes their deployment very challenging if not impossible. This difficulty stems from the unfounded implicit assumption that data-oriented routing policies directly reflect the underlying packet-level inter-domain policies. We find that to enable the more effective network utilization promised by data-oriented networking, essential caching incentives need to exist, and that data-oriented peering needs be considered separately from peering for packet forwarding.
Jarno Rajahalme, Mikko Särelä, Pekka Nikander, Sasu Tarkoma
CoNEXT4
2008 Windowing BitTorrent for Video-on-Demand: Not All is Lost with Tit-for-Tat
abstract
In this paper we present findings from our windowing BitTorrent simulations and show that by carefully optimizing other factors a reasonable level of performance can be achieved while leaving the original BitTorrent tit-for-tat mechanism intact. We compare the previously proposed windowing piece selection algorithms for BitTorrent and propose a new one, called the stretching window algorithm. We also propose a new method for reducing buffering times, adjusting the window size as the download progresses, and we show its effectiveness. We also observe that windowing BitTorrent algorithms exhibit steady state behavior, and that even a small level of altruism leads to significantly improved system performance.
Petri Savolainen, Niklas Raatikainen, Sasu Tarkoma
GLOBECOM3
2008 Dynamic filter merging and mergeability detection for publish/subscribe
Sasu Tarkoma
Pervasive Mob. Comput.1
2007 Context-Aware Learning for Intelligent Mobile Multimodal user Interfaces
abstract
The paper presents an association rule mining based learning approach for multimodal user interface adaptation in mobile environments. High-level knowledge about user preferences in multimodal interaction is inferred, using data mining techniques based on context parameters of the environment. The current approach facilitates automatic selection of multimodality capable interaction devices and their according rendering facilities for media output streams. An overview of the learning subsystem being part of the distributed communication sphere (DCS) management architecture, proposed within the EU IST-027617 project SPICE, will be introduced. Further the design of the learning approach will be discussed, including the definition and adaptation of snapshot data based on environment parameters. Frequent snapshots form the basis for learning, and therefore for the described association rule mining algorithm. Theoretical simulation results are presented and an outlook towards next research steps is given.
Ralf Kernchen, Klaus Moessner, Christian Räck, Oliver Sawade, Sasu Tarkoma, Stefan Arbanowski
PIMRC6
2007 Dynamic Filter Merging for Publish/Subscribe
abstract
Filter-based publish/subscribe and content-based routing have been proposed for flexible information dissemination in distributed environments. A content-based router is part of an overlay structure, in which each router forwards events to neighbouring routers and local clients based on their interests. Filter merging or summarization has been proposed as an optimization strategy in this environment. These techniques combine filters to reduce the number of propagated filters and thus the size of distributed state. In this paper, we present the algorithms for generic dynamic filter merging and discuss integration with routing tables. The algorithms are based on a formal framework of merging rules. Experimental results are examined and analyzed for both desktop systems and small devices. The results indicate that dynamic filter merging is feasible given that the workload is mergeable.
Sasu Tarkoma
WOWMOM1
2007 Spice: A Service Platform for Future Mobile IMS Services
abstract
Today's wireless and mobile service platforms are typically monolithic and centralized in nature, and they do not support heterogeneous service access and the sharing of service usage experience. New sources of revenue for providers are expected to include tailored, personalized, and dynamically composed services that are fast to market, cost efficient, and provide compelling user experience. To meet the current market needs, the SPICE service delivery platform extends the conventional IMS by supporting advanced added-value services that are composed of more primitive services. We describe the IMS role and functions in SPICE, and the use of ontology and Semantic Web technologies for achieving a better knowledge management in mobile service platforms.
Sasu Tarkoma, Ernö Kovacs, Herma Van Kranenburg, Erwin Postmann, Robert Seidl, Anna Fensel
WOWMOM1
2007 XML messaging for mobile devices: From requirements to implementation
Jaakko Kangasharju, Tancred Lindholm, Sasu Tarkoma
Comput. Networks3
2007 On the cost and safety of handoffs in content-based routing systems
Sasu Tarkoma, Jaakko Kangasharju
Comput. Networks1
2006 Fast and simple XML tree differencing by sequence alignment
abstract
With the advent of XML we have seen a renewed interest in methods for computing the difference between trees. Methods that include heuristic elements play an important role in practical applications due to the inherent complexity of the problem. We present a method for differencing XML as ordered trees based on mapping the problem to the domain of sequence alignment, applying simple and efficient heuristics in this domain, and transforming back to the tree domain. Our approach provides a method to quickly compute changes that are meaningful transformations on the XML tree level, and includes subtree move as a primitive operation. We evaluate the feasibility of our approach and benchmark it against a selection of existing differencing tools. The results show our approach to be feasible and to have the potential to perform on par with tools of a more complex design in terms of both output size and execution time.
Tancred Lindholm, Jaakko Kangasharju, Sasu Tarkoma
ACM Symposium on Document Engineering3
2006 TSR: Temporal Subspace Routing for Peer-to-Peer Data Sharing
abstract
In this paper we present the temporal subspace routing (TSR) technique for peer-to-peer environments that allows transparent exchange of information defined using metadata and queries based on user interests. The system unifies generic semantic matching, routing, caching, and access control. The technique supports continuous queries that are matched against metadata profiles of remote resources. Both queries and profiles are defined as subspaces of a multi-dimensional content space. Matched objects may be downloaded or synchronized. We present a generic data structure with optimizations for matching in this environment and discuss several use cases where the system may be applied. The mechanism utilizes the covering relation between queries and profiles. This allows automatic taxonomies of downloaded profiles and queries. Our main application is peer- to-peer and ad hoc metadata-based resource and file sharing.
Sasu Tarkoma
GLOBECOM1
2006 On Encrypting and Signing Binary XML Messages in the Wireless Environment
abstract
In the wireless world there has been much interest in alternate serialization formats for XML data, mostly driven by the weak capabilities of both devices and networks. However, an alternate serialization format is not easily made compatible with XML security features such as encryption and signing. We consider here ways to integrate an alternate format with security, and present a solution that we see as a viable alternative. In addition to this, we present extensive performance measurements, including ones on a mobile phone, on the effect of an alternate format when using XML-based security. These measurements indicate that, in the wireless world, reducing message sizes is the most pressing concern, and that processing efficiency gains of an alternate format are a much lesser concern
Jaakko Kangasharju, Tancred Lindholm, Sasu Tarkoma
ICWS3
2006 Fuego: Experiences with Mobile Data Communication and Synchronization
abstract
In this paper, we present a summary of our experiences with mobile middleware research in the four-year Fuego Core project. The presented work focuses on data communication and synchronization. We present three middleware services for data communication and synchronization, namely the messaging, event, and file synchronizer services, and discuss their development and usage. We conclude with an integrated architecture of these services and the lessons we have learned
Sasu Tarkoma, Jaakko Kangasharju, Tancred Lindholm, Kimmo E. E. Raatikainen
PIMRC1
2006 Optimizing content-based routers: posets and forests
Sasu Tarkoma, Jaakko Kangasharju
Distributed Comput.1
2005 Handover cost and mobility-safety of content streams
abstract
In this paper we investigate the handover cost and mobility-safety of content streams. Content streams are continuous flows of information from one node in a distributed network to another. The flows are established using publish/subscribe primitives and content-based routing of information. We examine two useful properties for mobility-aware content routing systems, namely completeness and mobility-safety. Then we determine the topology update cost for three interesting topologies, a number of optimizations, and show that if completeness cannot be assumed the signalling cost is considerably higher and content-based flooding needs to be used. We present simulation results for subscriber mobility for the investigated protocols. Both theoretical and experimental results show that rendezvous-points may be used to significantly reduce the signalling cost of handovers.
Sasu Tarkoma, Jaakko Kangasharju
MSWiM1
2005 Xebu: A Binary Format with Schema-Based Optimizations for XML Data
Jaakko Kangasharju, Sasu Tarkoma, Tancred Lindholm
WISE2