VLDB 2026 Research / reviewers in the wild / expert
Sara Khalifa
dblp:132/3883
· DBLP profile ↗
26ranked-venue papers
5as first author
12since 2021 · last 2026
0000-0002-3417-2834ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 5 since 2021Computer networks · 7 · 2 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 6 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 first-author · 1 since 2021Security and privacy · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SolarTrack: Exploring the Continuous Tracking Capabilities of Wearable Solar HarvestersabstractContinuous tracking is often thought to require specialised, actively powered sensors. Yet energy harvesters already embedded in commercial devices, such as Garmin solar-powered smartwatches, generate energy signals that inherently carry continuous variations linked to user motion and environment. Prior studies have shown that these signals are sufficient for classification tasks such as human activity and gesture recognition by exploiting class-distinguishing cues. However, whether they can support continuous trajectory tracking has remained an open question-until now.In this paper, we present the first fundamental study of continuous hand trajectory tracking with wearable solar harvesters. Through a novel radiometric model, we analytically link photovoltaic (PV) power to the solar cell’s geometric configuration, exposing both the promise of energy signals for tracking and their core limitations: the positional ambiguities and distortions that arise when power is used directly for positioning. To resolve this ambiguity, we propose SolarTrack, a framework that embeds a radiometric model as a physical backbone within a sequence-learning pipeline, enforcing cycle consistency between data-driven predictions and physical feasibility. This yields the first standalone solar-based tracker capable of estimating continuous hand motion directly from harvested energy signals.To validate this, we built a wearable prototype with a solar panel and IMU and collected the first dataset pairing motion-capture ground truth with harvested energy from 15 participants (700k samples). Results show that solar signals alone achieve subdecimeter tracking accuracy, outperforming purely data-driven baselines and only 1.6 cm worse compared to the IMU tracker despite having access to only 1D power signal. Furthermore, when fused with IMU, it further boosts IMU performance by 13%. Yasien Ghalwash, Abdelwahed Khamis, Muhammad Moid Sandhu, Sara Khalifa, Raja Jurdak |
PerCom | 4 |
| 2025 | NeuralPrefix: A Zero-shot Sensory Data Imputation PluginabstractReal-world sensing challenges such as sensor failures, communication issues, and power constraints lead to data intermittency. An issue that is known to undermine the traditional classification task that assumes a continuous data stream. Previous works addressed this issue by designing bespoke solutions (i.e. task-specific and/or modality-specific imputation). These approaches, while effective for their intended purposes, had limitations in their applicability across different tasks and sensor modalities. This raises an important question: Can we build a task-agnostic imputation pipeline that is transferable to new sensors without requiring additional training? In this work, we formalise the concept of zero-shot imputation and propose a novel approach that enables the adaptation of pre-trained models to handle data intermittency. This framework, named NeuralPrefix, is a generative neural component that precedes a task model during inference, filling in gaps caused by data intermittency. NeuralPrefix is built as a continuous dynamical system, where its internal state can be estimated at any point in time by solving an Ordinary Differential Equation (ODE). This approach allows for a more versatile and adaptable imputation method, overcoming the limitations of task-specific and modality-specific solutions. We conduct a comprehensive evaluation of NeuralPrefix on multiple sensory datasets, demonstrating its effectiveness across various domains. When tested on intermittent data with a high 50% missing data rate, NeuralPreifx accurately recovers all the missing samples, achieving SSIM score between 0.93-0.96. Zero-shot evaluations show that NeuralPrefix generalises well to unseen datasets, even when the measurements come from a different modality. Project page: https://neuralprefix.github.io/ Abdelwahed Khamis, Sara Khalifa |
PerCom | 2 |
| 2025 | VEH-Attack: Stealthy Tracking of Train Passengers With Side-Channel Attack on Vibration Energy Harvesting WearablesabstractVibration energy harvesting (VEH) has emerged as a viable option for mobile devices that serves the dual purpose of generating power and sensing ambient vibrations. This paper highlights the location privacy leakage resulting from unrestricted access to seemingly innocuous VEH data on mobile devices. We present VEH-Attack, a side-channel attack that exploits an inference model and VEH data patterns generated from train vibrations, enabling precise tracking of train passengers. VEH-Attack achieves an accuracy of 97% and 83.13% for VEH derived data and actual VEH data, respectively, for trip length of 6 stations with the accuracy reaching 100% for longer trip lengths. Marzieh Jalal Abadi, Sara Khalifa, Mahbub Hassan, Salil S. Kanhere, Mohamed Ali Kâafar |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2023 | Survey of Deep Representation Learning for Speech Emotion RecognitionabstractTraditionally, speech emotion recognition (SER) research has relied on manually handcrafted acoustic features using feature engineering. However, the design of handcrafted features for complex SER tasks requires significant manual effort, which impedes generalisability and slows the pace of innovation. This has motivated the adoption of representation learning techniques that can automatically learn an intermediate representation of the input signal without any manual feature engineering. Representation learning has led to improved SER performance and enabled rapid innovation. Its effectiveness has further increased with advances in deep learning (DL), which has facilitateddeep representation learningwhere hierarchical representations are automatically learned in a data-driven manner. This article presents the first comprehensive survey on the important topic of deep representation learning for SER. We highlight various techniques, related challenges and identify important future areas of research. Our survey bridges the gap in the literature since existing surveys either focus on SER with hand-engineered features or representation learning in the general setting without focusing on SER. Siddique Latif, Rajib Rana, Sara Khalifa, Raja Jurdak, Junaid Qadir 0001, Björn W. Schuller |
IEEE Trans. Affect. Comput. | 3 |
| 2023 | Self Supervised Adversarial Domain Adaptation for Cross-Corpus and Cross-Language Speech Emotion RecognitionabstractDespite the recent advancement in speech emotion recognition (SER) within a single corpus setting, the performance of these SER systems degrades significantly for cross-corpus and cross-language scenarios. The key reason is the lack of generalisation in SER systems towards unseen conditions, which causes them to perform poorly in cross-corpus and cross-language settings. Recent studies focus on utilising adversarial methods to learn domain generalised representation for improving cross-corpus and cross-language SER to address this issue. However, many of these methods only focus on cross-corpus SER without addressing the cross-language SER performance degradation due to a larger domain gap between source and target language data. This contribution proposes an adversarial dual discriminator (ADDi) network that uses the three-players adversarial game to learn generalised representations without requiring any target data labels. We also introduce a self-supervised ADDi (sADDi) network that utilises self-supervised pre-training with unlabelled data. We propose synthetic data generation as a pretext task in sADDi, enabling the network to produce emotionally discriminative and domain invariant representations and providing complementary synthetic data to augment the system. The proposed model is rigorously evaluated using five publicly available datasets in three languages and compared with multiple studies on cross-corpus and cross-language SER. Experimental results demonstrate that the proposed model achieves improved performance compared to the state-of-the-art methods. Siddique Latif, Rajib Rana, Sara Khalifa, Raja Jurdak, Björn W. Schuller |
IEEE Trans. Affect. Comput. | 3 |
| 2023 | Multitask Learning From Augmented Auxiliary Data for Improving Speech Emotion RecognitionabstractDespite the recent progress in speech emotion recognition (SER), state-of-the-art systems lack generalisation across different conditions. A key underlying reason for poor generalisation is the scarcity of emotion datasets, which is a significant roadblock to designing robust machine learning (ML) models. Recent works in SER focus on utilising multitask learning (MTL) methods to improve generalisation by learning shared representations. However, most of these studies propose MTL solutions with the requirement of meta labels for auxiliary tasks, which limits the training of SER systems. This paper proposes an MTL framework (MTL-AUG) that learns generalised representations from augmented data. We utilise augmentation-type classification and unsupervised reconstruction as auxiliary tasks, which allow training SER systems on augmented data without requiring any meta labels for auxiliary tasks. The semi-supervised nature of MTL-AUG allows for the exploitation of the abundant unlabelled data to further boost the performance of SER. We comprehensively evaluate the proposed framework in the following settings: (1) within corpus, (2) cross-corpus and cross-language, (3) noisy speech, (4) and adversarial attacks. Our evaluations using the widely used IEMOCAP, MSP-IMPROV, and EMODB datasets show improved results compared to existing state-of-the-art methods. Siddique Latif, Rajib Rana, Sara Khalifa, Raja Jurdak, Björn W. Schuller |
IEEE Trans. Affect. Comput. | 3 |
| 2022 | Distilling Representational Similarity using Centered Kernel Alignment (CKA)
Aninda Saha, Alina Bialkowski, Sara Khalifa |
BMVC | 3 |
| 2022 | Demo Abstract: A Real Time Control System for Replaying Motion DataabstractThis paper demonstrates a laboratory setup to replay previously collected experimental human motion data in real-time, allowing data re-purposing for activity recognition. The setup proposed is to examine the suitability of simulating motion by replaying a previ-ously collected data using an Electrodynamic Shaker - a table that can generate motion based on pre-recorded accelerometer traces. Building upon the information gathered from the sensor input, ac-curate continuous motion data from human dynamic movements can be recreated. By reusing data from past projects with novel or alternative sensor attachments, new datasets are created gen-erating activity data for movement recognition without requiring volunteers or time consuming data collection. This demo illustrates the potential for accurate low-cost in-house motion data collection without requiring field experiments. Margaux Edwards, Sara Khalifa |
IPSN | 2 |
| 2022 | Poster Abstract: Trade-off Analysis of Inference Accuracy and Resource Usage for Energy-Positive Activity RecognitionabstractEnergy-positive activity recognition classifies human activities, including walking, running, and sitting, while harvesting kinetic energy from such activities. In this setting, the device's lifetime de-pends on the user's activity profile and the resources needed to run inference to classify activities. Thus, the selection of machine learning classification models for energy-positive activity recognition must consider both model's classification accuracy and energy con-sumption compared to the harvested energy from human activities. In this paper, we study the trade-off between accuracy and resource usage of a neural network model when different feature extraction techniques are used. Our results indicate that an on-board sched-uling algorithm can be used to dynamically switch between the optimal feature input tuned for accuracy and energy consumption. Minh Tuan Tran, Muhammad Moid Sandhu, Sara Khalifa, Gowri Sankar Ramachandran, Raja Jurdak |
IPSN | 3 |
| 2022 | Multi-Task Semi-Supervised Adversarial Autoencoding for Speech Emotion RecognitionabstractInspite the emerging importance of Speech Emotion Recognition (SER), the state-of-the-art accuracy is quite low and needs improvement to make commercial applications of SER viable. A key underlying reason for the low accuracy is the scarcity of emotion datasets, which is a challenge for developing any robust machine learning model in general. In this article, we propose a solution to this problem: a multi-task learning framework that uses auxiliary tasks for which data is abundantly available. We show that utilisation of this additional data can improve the primary task of SER for which only limited labelled data is available. In particular, we use gender identifications and speaker recognition as auxiliary tasks, which allow the use of very large datasets, e. g., speaker classification datasets. To maximise the benefit of multi-task learning, we further use an adversarial autoencoder (AAE) within our framework, which has a strong capability to learn powerful and discriminative features. Furthermore, the unsupervised AAE in combination with the supervised classification networks enables semi-supervised learning which incorporates a discriminative component in the AAE unsupervised training pipeline. This semi-supervised learning essentially helps to improve generalisation of our framework and thus leads to improvements in SER performance. The proposed model is rigorously evaluated for categorical and dimensional emotion, and cross-corpus scenarios. Experimental results demonstrate that the proposed model achieves state-of-the-art performance on two publicly available datasets. Siddique Latif, Rajib Rana, Sara Khalifa, Raja Jurdak, Julien Epps, Björn W. Schuller |
IEEE Trans. Affect. Comput. | 3 |
| 2021 | SolAR: Energy Positive Human Activity Recognition using Solar CellsabstractThe high power consumption of inertial activity sensors limits the battery lifetime of today's wearable devices. Recent studies promise to extend the lifetime of wearable devices by translating kinetic energy from human movements into electrical energy while using the harvesting signal to replace conventional activity sensors. However, in human-centric applications, the amount of harvested kinetic energy is not enough to power a real-time activity recognition algorithm and run the wearable device perpetually. In this paper, we propose Solar based human Activity Recognition (SolAR), which uses solar cells simultaneously as an activity sensor as well as an energy source. Our key observation is that the power available from a wrist-worn solar cell changes dynamically while a person moves, encoding information about the underlying activity. We collect empirical solar energy data to explore its activity sensing potential and implement the activity recognition pipeline on an ultra low-power micro-controller unit to evaluate the end-to-end power consumption of the system. Our analysis reveals that SolAR improves activity recognition accuracy by up to 8.3% and harvests more than one order of magnitude higher power compared to its kinetic counterpart. This enables SolAR to generate more energy than required for the entire activity recognition pipeline, which we term as energy positive activity recognition, achieving uninterrupted, autonomous, self-powered and real-time operation. Muhammad Moid Sandhu, Sara Khalifa, Kai Geissdoerfer, Raja Jurdak, Marius Portmann |
PerCom | 2 |
| 2021 | Task Scheduling for Energy-Harvesting-Based IoT: A Survey and Critical AnalysisabstractThe Internet of Things (IoT) has important applications in our daily lives, including health and fitness tracking, environmental monitoring, and transportation. However, sensor nodes in IoT suffer from the limited lifetime of batteries resulting from their finite energy availability. A promising solution is to harvest energy from environmental sources, such as solar, kinetic, thermal, and radio-frequency (RF) waves, for perpetual and continuous operation of IoT sensor nodes. In addition to energy generation, recently energy harvesters have been used for context detection, eliminating the need for conventional activity sensors (e.g., accelerometers), saving space, cost, and energy consumption. Using energy harvesters for simultaneous sensing and energy harvesting enables energy positive sensing-an important and emerging class of sensors, which harvest more energy than required for context detection and the additional energy can be used to power other components of the system. Although simultaneous sensing and energy harvesting is an important step forward toward autonomous self-powered sensor nodes, the energy and information availability can be still intermittent, unpredictable, and temporally misaligned with various computational tasks on the sensor node. This article provides a comprehensive survey on task scheduling algorithms for the emerging class of energy harvesting-based sensors (i.e., energy positive sensors) to achieve the sustainable operation of IoT. We discuss inherent differences between conventional sensing and energy positive sensing and provide an extensive critical analysis for devising revised task scheduling algorithms incorporating this new class of sensors. Finally, we outline future research directions toward the implementation of autonomous and self-powered IoT. Muhammad Moid Sandhu, Sara Khalifa, Raja Jurdak, Marius Portmann |
IEEE Internet Things J. | 2 |
| 2020 | Augmenting Generative Adversarial Networks for Speech Emotion RecognitionabstractGenerative adversarial networks (GANs) have shown potential in learning emotional attributes and generating new data samples. However, their performance is usually hindered by the unavailability of larger speech emotion recognition (SER) data. In this work, we propose a framework that utilises the mixup data augmentation scheme to augment the GAN in feature learning and generation. To show the effectiveness of the proposed framework, we present results for SER on (i) synthetic feature vectors, (ii) augmentation of the training data with synthetic features, (iii) encoded features in compressed representation. Our results show that the proposed framework can effectively learn compressed emotional representations as well as it can generate synthetic samples that help improve performance in within-corpus and cross-corpus evaluation. Siddique Latif, Muhammad Asim 0005, Rajib Rana, Sara Khalifa, Raja Jurdak, Björn W. Schuller |
INTERSPEECH | 4 |
| 2020 | Deep Architecture Enhancing Robustness to Noise, Adversarial Attacks, and Cross-Corpus Setting for Speech Emotion RecognitionabstractSpeech emotion recognition systems (SER) can achieve high accuracy when the training and test data are identically distributed, but this assumption is frequently violated in practice and the performance of SER systems plummet against unforeseen data shifts. The design of robust models for accurate SER is challenging, which limits its use in practical applications. In this paper we propose a deeper neural network architecture wherein we fuse DenseNet, LSTM and Highway Network to learn powerful discriminative features which are robust to noise. We also propose data augmentation with our network architecture to further improve the robustness. We comprehensively evaluate the architecture coupled with data augmentation against (1) noise, (2) adversarial attacks and (3) cross-corpus settings. Our evaluations on the widely used IEMOCAP and MSP-IMPROV datasets show promising results when compared with existing studies and state-of-the-art models. Siddique Latif, Rajib Rana, Sara Khalifa, Raja Jurdak, Björn W. Schuller |
INTERSPEECH | 3 |
| 2020 | Poster Abstract: Federated Learning for Speech Emotion Recognition ApplicationsabstractPrivacy concerns are considered one of the major challenges in the applications of speech emotion recognition (SER) as it involves the complete sharing of speech data, which can bring threatening consequences to people’s lives. Federated learning is an effective technique to avoid privacy infringement by involving multiple participants to collaboratively learn a shared model without revealing their local data. In this work, we evaluated federated learning for SER using a publicly available dataset. Our preliminary results show that speech emotion recognition can benefit from federated learning by not exporting sensitive user data to central servers, while achieving promising results compared to the state-of-the-art. Siddique Latif, Sara Khalifa, Rajib Rana, Raja Jurdak |
IPSN | 2 |
| 2020 | Towards Energy Positive Sensing using Kinetic Energy HarvestersabstractConventional systems for motion context detection rely on batteries to provide the energy required for sampling a motion sensor. Batteries, however, have limited capacity and, once depleted, have to be replaced or recharged. Kinetic Energy Harvesting (KEH) allows to convert ambient motion and vibration into usable electricity and can enable batteryless, maintenance free operation of motion sensors. The signal from a KEH transducer correlates with the underlying motion and may thus directly be used for context detection, saving space, cost and energy by omitting the accelerometer. Previous work uses the open circuit or the capacitor voltage for sensing without using the harvested energy to power a load. In this paper, we propose to use other sensing points in the KEH circuit that offer information-rich sensing signals while the energy from the harvester is used to power a load. We systematically analyze multiple sensing signals available in different KEH architectures and compare their performance in a transport mode detection case study. To this end, we develop four hardware prototypes, conduct an extensive measurement campaign and use the data to train and evaluate different classifiers. We show that sensing the harvesting current signal from a transducer can be energy positive, delivering up to ten times as much power as it consumes for signal acquisition, while offering comparable detection accuracy to the accelerometer signal for most of the considered transport modes. Muhammad Moid Sandhu, Kai Geissdoerfer, Sara Khalifa, Raja Jurdak, Marius Portmann, Branislav Kusy |
PerCom | 3 |
| 2020 | EnTrans: Leveraging Kinetic Energy Harvesting Signal for Transportation Mode DetectionabstractMonitoring the daily transportation modes of an individual provides useful information in many application domains, such as urban design, real-time journey recommendation, and providing location-based services. In existing systems, accelerometer and GPS are the dominantly used signal sources for transportation context monitoring which drain out the limited battery life of the wearable devices very quickly. To resolve the high energy consumption issue, in this paper, we present EnTrans, which enables transportation mode detection by using only the kinetic energy harvester as an energy-efficient signal source. The proposed idea is based on the intuition that the vibrations experienced by the passenger during traveling with different transportation modes are distinctive. Thus, voltage signal generated by the energy harvesting devices should contain sufficient features to distinguish different transportation modes. We evaluate our system using over 28 h of data, which is collected by eight individuals using a practical energy harvesting prototype. The evaluation results demonstrate that EnTrans is able to achieve an overall accuracy over 92% in classifying five different modes while saving more than 34% of the system power compared to conventional accelerometer-based approaches. Guohao Lan, Weitao Xu, Dong Ma 0001, Sara Khalifa, Mahbub Hassan, Wen Hu 0001 |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2019 | Direct Modelling of Speech Emotion from Raw SpeechabstractSpeech emotion recognition is a challenging task and heavily depends on hand-engineered acoustic features, which are typically crafted to echo human perception of speech signals. However, a filter bank that is designed from perceptual evidence is not always guaranteed to be the best in a statistical modelling framework where the end goal is for example emotion classification. This has fuelled the emerging trend of learning representations from raw speech especially using deep learning neural networks. In particular, a combination of Convolution Neural Networks (CNNs) and Long Short Term Memory (LSTM) have gained great traction for the intrinsic property of LSTM in learning contextual information crucial for emotion recognition; and CNNs been used for its ability to overcome the scalability problem of regular neural networks. In this paper, we show that there are still opportunities to improve the performance of emotion recognition from the raw speech by exploiting the properties of CNN in modelling contextual information. We propose the use of parallel convolutional layers to harness multiple temporal resolutions in the feature extraction block that is jointly trained with the LSTM based classification network for the emotion recognition task. Our results suggest that the proposed model can reach the performance of CNN trained with hand-engineered features from both IEMOCAP and MSP-IMPROV datasets. Siddique Latif, Rajib Rana, Sara Khalifa, Raja Jurdak, Julien Epps |
INTERSPEECH | 3 |
| 2019 | KEH-Gait: Using Kinetic Energy Harvesting for Gait-based User Authentication SystemsabstractWith the rapid development of sensor networks and embedded computing technologies, miniaturized wearable healthcare monitoring devices have become practically feasible. For many of these devices, accelerometer-based user authentication systems by gait analysis are becoming a hot research topic. However, a major bottleneck of such system is it requires continuous sampling of accelerometer, which reduces battery life of wearable sensors. In this paper, we present KEH-Gait, which advocates use of output voltage signal from kinetic energy harvester (KEH) as the source for gait recognition. KEH-Gait is motivated by the prospect of significant power saving by not having to sample the accelerometer at all. Indeed, our measurements show that, compared to conventional accelerometer-based gait detection, KEH-Gait can reduce energy consumption by 82.15 percent. The feasibility of KEH-Gait is based on the fact that human gait has distinctive movement patterns for different individuals, which is expected to leave distinctive patterns for KEH as well. We evaluate the performance of KEH-Gait using two different types of KEH hardware on a data set of 20 subjects. Our experiments demonstrate that, although KEH-Gait yields slightly lower accuracy than accelerometer-based gait detection when single step is used, the accuracy problem can be overcome by the proposed Probability-based Multi-Step Sparse Representation Classification (PMSSRC). Moreover, the security analysis shows that the EER of KEH-Gait against an active spoofing attacker is 11.2 and 14.1 percent using two different types of KEH hardware, respectively. Weitao Xu, Guohao Lan, Sara Khalifa, Mahbub Hassan, Neil W. Bergmann, Wen Hu 0001 |
IEEE Trans. Mob. Comput. | 4 |
| 2018 | HARKE: Human Activity Recognition from Kinetic Energy Harvesting Data in Wearable DevicesabstractKinetic energy harvesting (KEH) may help combat battery issues in wearable devices. While the primary objective of KEH is to generate energy from human activities, the harvested energy itself contains information about human activities that most wearable devices try to detect using motion sensors. In principle, it is therefore possible to use KEH both as a power generator and a sensor for human activity recognition (HAR), saving sensor-related power consumption. Our aim is to quantify the potential of human activity recognition from kinetic energy harvesting (HARKE). We evaluate the performance of HARKE using two independent datasets: (i) a public accelerometer dataset converted into KEH data through theoretical modeling; and (ii) a real KEH dataset collected from volunteers performing activities of daily living while wearing a data-logger that we built of a piezoelectric energy harvester. Our results show that HARKE achieves an accuracy of 80 to 95 percent, depending on the dataset and the placement of the device on the human body. We conduct detailed power consumption measurements to understand and quantify the power saving opportunity of HARKE. The results demonstrate that HARKE can save 79 percent of the overall system power consumption of conventional accelerometer-based HAR. Sara Khalifa, Guohao Lan, Mahbub Hassan, Aruna Seneviratne, Sajal K. Das 0001 |
IEEE Trans. Mob. Comput. | 1 |
| 2017 | KEH-Gait: Towards a Mobile Healthcare User Authentication System by Kinetic Energy Harvesting
Weitao Xu, Guohao Lan, Sara Khalifa, Neil W. Bergmann, Mahbub Hassan, Wen Hu 0001 |
NDSS | 4 |
| 2017 | VEH-COM: Demodulating vibration energy harvesting for short range communicationabstractThis paper investigates the possibility of using a vibration energy harvesting (VEH) device as a communication receiver. By modulating the ambient vibration energy using a transmitting speaker, and demodulating the harvested power at the receiving VEH, we aim to transmit small amounts of data at low rates between two proximate devices. The key advantage of using VEH as a receiver is that the modulated sound waves can be successfully demodulated directly from the harvested power without employing the power-consuming digital signal processing (DSP), which makes a VEH receiver significantly more power efficient than a conventional microphone-based decoder. To address the extremely narrow bandwidth of VEH, we design a simple ON-OFF keying modulation, but optimized for VEH hardware. Experiments with a real VEH device shows that, at a distance of 2 cm, a laptop speaker with the proposed modulation scheme can achieve 30 bps communication for a target bit error rate of less than 1%, which would enable many emerging short range applications, such as mobile payment. The communication range of a laptop can be extended to 80 cm for 5 bps, allowing a range of other audio-based device-to-device communications, such as a web advertisement on a laptop browser transferring tokens to a nearby smartphone. We also demonstrate that the proposed VEH-based sound decoding is resilient to background noise, thanks to its extremely narrow power harvesting bandwidth, which works as a natural noise filter. Guohao Lan, Weitao Xu, Sara Khalifa, Mahbub Hassan, Wen Hu 0001 |
PerCom | 3 |
| 2016 | Feasibility and accuracy of hotword detection using vibration energy harvesterabstractVibration energy harvesting (VEH) is a promising source of renewable energy that can be used to extend battery life of next generation mobile devices. In this paper, we study the feasibility and accuracy of VEH for detecting hotwords, such as “OK Google”, used by popular voice control applications to distinguish user commands from other conversations. The idea of using power signals of VEH to detect hotwords is based on the fact that human voice creates vibrations in the air, which could be potentially picked up by the VEH hardware inside a mobile device. Using off-the-shelf VEH product, we conduct a comprehensive experimental study involving 8 subjects. We analyse two possible usage scenarios for the VEH hardware. In the first scenario, the user is not required to talk directly to the device (indirect), but the VEH is expected to pick up the ambient vibrations caused by user-generated sound waves. In the second, the user is expected to direct his voice to the VEH (direct) and talk to it from a close distance. For both usage scenarios, we evaluate two types of hotword detection, speaker-independent and speaker-dependent. We find that VEH can detect hotwords more accurately in the direct scenario compared to the indirect. For the direct scenario, our results show that a simple Decision Tree classifier can detect hotwords from VEH signals with accuracies of 73% and 85%, respectively, for speaker-independent and speaker-dependent detections. Finally, we show that these accuracies are comparable to what could be achieved with an accelerometer sampled at 200 Hz. Sara Khalifa, Mahbub Hassan, Aruna Seneviratne |
WoWMoM | 1 |
| 2015 | Pervasive self-powered human activity recognition without the accelerometerabstractConventional human activity recognition (HAR) relies on accelerometers to frequently sample human motion (acceleration). Unfortunately, power consumption of accelerometers becomes a bottleneck for realising pervasive self-powering HAR as the amount of power that can be practically harvested from the environment is very small. Instead of using accelerometer, this paper advocates the use of energy harvesting power signal as the source of HAR when motion (kinetic) energy is being harvested to power the device. The proposed use of harvested power for classifying human activities is motivated by the fact that different activities produce kinetic energy in a different way leaving their signatures in the harvested power signal. Using information theoretic analysis of experimental data, we show that many standard statistical features provide significant information gain when the kinetic power signal is used for discriminating between different activities, confirming its potential use for HAR. We have evaluated activity recognition accuracy for kinetic power signal based HAR using 14 different sets of common activities each containing between 2-10 different activities to be classified. HAR accuracies varied between 68% to 100% depending on the set of activities. The average accuracy over all activity sets is 83%, which is within 13% of what could be achieved with an accelerometer without any power constraints. Sara Khalifa, Mahbub Hassan, Aruna Seneviratne |
PerCom | 1 |
| 2013 | Adaptive pedestrian activity classification for indoor dead reckoning systemsabstractA pedestrian activity classification (PAC) system classifies pedestrian motion data into activities related to the usage of specific building facilities, such as going up on an escalator or descending a staircase. Recent studies confirm that use of PAC significantly reduces indoor localization errors of a pedestrian dead reckoning (PDR) system as exact facility locations in the building can be retrieved from the floor map. However, classification complexity may become an issue for resource constraint mobile devices. We propose a novel PAC system that, instead of using a single complex classifier based on a large set of features, employs multiple simple classifiers each trained to classify only a subset of the activities using a small number of features. As the pedestrian moves around inside a building, the proposed adaptive-PAC dynamically switches to the right (simple) classifier based on the facilities that exist within the immediate proximity. By always using a simple classifier, adaptive-PAC has the potential to drastically reduce the average classification complexity for PAC-aided PDR systems. Using experimental data, we quantify and compare the performance of the proposed adaptive-PAC against the conventional PAC. We find that for typical shopping centers, adaptive-PAC reduces classification complexity by 91-97% without any degradation in classification accuracy rates. Sara Khalifa, Mahbub Hassan, Aruna Seneviratne |
IPIN | 1 |
| 2012 | Evaluating mismatch probability of activity-based map matching in indoor positioningabstractIf users are known to perform specific activities at specific locations within a building, then indoor positioning could be achieved by monitoring user activities and matching them to specific locations in a preloaded floor map. This is the fundamental idea behind activity-based map matching (AMM). For example, the user's smartphone could use the accelerometer readings to detect whether a user is using an escalator, and then match the current location of the user to the nearest escalator. AMM therefore could be used for frequently recalibrating location estimators to ground-truth values. This is especially useful for recalibrating pedestrian dead reckoning (PDR), which can estimate indoor position if started from a known location, but error grows unboundedly with time or distance traveled. However, AMM is not perfect and could potentially cause mismatches by matching the current location of the user to a wrong location. In this paper we propose a methodology and derive a closed-form expression for mismatch probability as a function of PDR sensor error and proximity between two facilities. By applying our methodology to a practical indoor complex (Sydney airport) we find several interesting results: (1) that mismatch probability is spatially non-uniform, i.e., it can be different in different parts of the floor, (2) for some specific facilities, mismatch probability can be very high (up to 80%), and (3) if escalators could be distinguished from lifts with high accuracy, we could reduce mismatch probability significantly (by up to 68%). Sara Khalifa, Mahbub Hassan |
IPIN | 1 |