VLDB 2026 Research / reviewers in the wild / expert
Rajib Rana
dblp:136/5943 · also Rajib Kumar Rana
· DBLP profile ↗
30ranked-venue papers
6as first author
11since 2021 · last 2026
0000-0002-0506-2409ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 16 · 1 first-author · 10 since 2021Computer networks · 9 · 2 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6Databases, data management, data science and information retrieval · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FedMLAC: Mutual learning driven heterogeneous federated audio classificationabstractFederated Learning (FL) offers a privacy-preserving framework for training audio classification (AC) models across decentralized clients without sharing raw data. However, Federated Audio Classification faces three major challenges: data heterogeneity , model heterogeneity , and data corruption , which degrade performance in real-world settings. While existing methods often address these issues separately, a unified solution remains underexplored. We propose FedMLAC, a mutual learning-based FL framework that tackles all three challenges simultaneously. Each client maintains a personalized local AC model and a lightweight, globally shared Plug-in model. These models interact via bidirectional knowledge distillation, enabling global knowledge sharing while adapting to local data distributions, thus addressing both data and model heterogeneity. To counter data corruption, we introduce a Layer-wise Pruning Aggregation (LPA) strategy that filters anomalous Plug-in updates based on parameter deviations during aggregation. Extensive experiments on four diverse AC benchmarks, including both speech and non-speech tasks, show that FedMLAC consistently outperforms state-of-the-art baselines in classification accuracy and robustness to noisy data. Rajib Rana, Di Wu 0050, Youyang Qu, Xiaohui Tao 0001, Ji Zhang 0001, Carlos Busso, Palaiahnakote Shivakumara |
Pattern Recognit. | 2 |
| 2025 | Multitask Transformer for Cross-Corpus Speech Emotion RecognitionabstractDeep learning has significantly advanced the field of Speech Emotion Recognition (SER), yet its efficacy in cross-corpus scenarios remains a challenge. To overcome this limitation, recent studies demonstrate the success of multitask learning, which uses auxiliary tasks to reduce difference between source and target dataset (or transfer knowledge from source to target datasets). Despite the efforts, the overall accuracy for cross-corpus SER is still relatively low and needs attention. To improve performance, we propose a multitask framework with SER as the primary task and contrastive learning and information maximization as auxiliary tasks. We design the auxiliary tasks innovatively to use the target data without emotional labels to develop a better understanding of the target data. The core of our multitask framework is a pre-trained transformer. While transformers have gained attention in SER, their application to cross-corpus scenarios is still limited. Multimodal approaches for cross-corpus scenario is substantially limited as well. We use text as the second modality, developing separate multitask transformers for audio and text and conduct decision-level fusion during inference. We use publicly available and widely used speech corpora, including the IEMOCAP, MSP-IMPROV and EMO-DB databases. The results demonstrate the benefits of the proposed approach, achieving improved performance on the benchmark databases in cross-corpus settings. Chung Soo Ahn, Rajib Rana, Carlos Busso, Jagath C. Rajapakse |
IEEE Trans. Affect. Comput. | 2 |
| 2023 | Improving the Domain Adaptation of Retrieval Augmented Generation (RAG) Models for Open Domain Question AnsweringabstractAbstract Retrieval Augment Generation (RAG) is a recent advancement in Open-Domain Question Answering (ODQA). RAG has only been trained and explored with a Wikipedia-based external knowledge base and is not optimized for use in other specialized domains such as healthcare and news. In this paper, we evaluate the impact of joint training of the retriever and generator components of RAG for the task of domain adaptation in ODQA. We propose RAG-end2end, an extension to RAG that can adapt to a domain-specific knowledge base by updating all components of the external knowledge base during training. In addition, we introduce an auxiliary training signal to inject more domain-specific knowledge. This auxiliary signal forces RAG-end2end to reconstruct a given sentence by accessing the relevant information from the external knowledge base. Our novel contribution is that, unlike RAG, RAG-end2end does joint training of the retriever and generator for the end QA task and domain adaptation. We evaluate our approach with datasets from three domains: COVID-19, News, and Conversations, and achieve significant performance improvements compared to the original RAG model. Our work has been open-sourced through the HuggingFace Transformers library, attesting to our work’s credibility and technical consistency. Shamane Siriwardhana, Rivindu Weerasekera, Tharindu Kaluarachchi, Elliott Wen, Rajib Rana, Suranga Nanayakkara |
Trans. Assoc. Comput. Linguistics | 5 |
| 2023 | Survey of Deep Representation Learning for Speech Emotion RecognitionabstractTraditionally, speech emotion recognition (SER) research has relied on manually handcrafted acoustic features using feature engineering. However, the design of handcrafted features for complex SER tasks requires significant manual effort, which impedes generalisability and slows the pace of innovation. This has motivated the adoption of representation learning techniques that can automatically learn an intermediate representation of the input signal without any manual feature engineering. Representation learning has led to improved SER performance and enabled rapid innovation. Its effectiveness has further increased with advances in deep learning (DL), which has facilitateddeep representation learningwhere hierarchical representations are automatically learned in a data-driven manner. This article presents the first comprehensive survey on the important topic of deep representation learning for SER. We highlight various techniques, related challenges and identify important future areas of research. Our survey bridges the gap in the literature since existing surveys either focus on SER with hand-engineered features or representation learning in the general setting without focusing on SER. Siddique Latif, Rajib Rana, Sara Khalifa, Raja Jurdak, Junaid Qadir 0001, Björn W. Schuller |
IEEE Trans. Affect. Comput. | 2 |
| 2023 | Self Supervised Adversarial Domain Adaptation for Cross-Corpus and Cross-Language Speech Emotion RecognitionabstractDespite the recent advancement in speech emotion recognition (SER) within a single corpus setting, the performance of these SER systems degrades significantly for cross-corpus and cross-language scenarios. The key reason is the lack of generalisation in SER systems towards unseen conditions, which causes them to perform poorly in cross-corpus and cross-language settings. Recent studies focus on utilising adversarial methods to learn domain generalised representation for improving cross-corpus and cross-language SER to address this issue. However, many of these methods only focus on cross-corpus SER without addressing the cross-language SER performance degradation due to a larger domain gap between source and target language data. This contribution proposes an adversarial dual discriminator (ADDi) network that uses the three-players adversarial game to learn generalised representations without requiring any target data labels. We also introduce a self-supervised ADDi (sADDi) network that utilises self-supervised pre-training with unlabelled data. We propose synthetic data generation as a pretext task in sADDi, enabling the network to produce emotionally discriminative and domain invariant representations and providing complementary synthetic data to augment the system. The proposed model is rigorously evaluated using five publicly available datasets in three languages and compared with multiple studies on cross-corpus and cross-language SER. Experimental results demonstrate that the proposed model achieves improved performance compared to the state-of-the-art methods. Siddique Latif, Rajib Rana, Sara Khalifa, Raja Jurdak, Björn W. Schuller |
IEEE Trans. Affect. Comput. | 2 |
| 2023 | Multitask Learning From Augmented Auxiliary Data for Improving Speech Emotion RecognitionabstractDespite the recent progress in speech emotion recognition (SER), state-of-the-art systems lack generalisation across different conditions. A key underlying reason for poor generalisation is the scarcity of emotion datasets, which is a significant roadblock to designing robust machine learning (ML) models. Recent works in SER focus on utilising multitask learning (MTL) methods to improve generalisation by learning shared representations. However, most of these studies propose MTL solutions with the requirement of meta labels for auxiliary tasks, which limits the training of SER systems. This paper proposes an MTL framework (MTL-AUG) that learns generalised representations from augmented data. We utilise augmentation-type classification and unsupervised reconstruction as auxiliary tasks, which allow training SER systems on augmented data without requiring any meta labels for auxiliary tasks. The semi-supervised nature of MTL-AUG allows for the exploitation of the abundant unlabelled data to further boost the performance of SER. We comprehensively evaluate the proposed framework in the following settings: (1) within corpus, (2) cross-corpus and cross-language, (3) noisy speech, (4) and adversarial attacks. Our evaluations using the widely used IEMOCAP, MSP-IMPROV, and EMODB datasets show improved results compared to existing state-of-the-art methods. Siddique Latif, Rajib Rana, Sara Khalifa, Raja Jurdak, Björn W. Schuller |
IEEE Trans. Affect. Comput. | 2 |
| 2023 | Emotion Intensity and its Control for Emotional Voice ConversionabstractEmotional voice conversion (EVC) seeks to convert the emotional state of an utterance while preserving the linguistic content and speaker identity. In EVC, emotions are usually treated as discrete categories overlooking the fact that speech also conveys emotions with various intensity levels that the listener can perceive. In this paper, we aim to explicitly characterize and control the intensity of emotion. We propose to disentangle the speaker style from linguistic content and encode the speaker style into a style embedding in a continuous space that forms the prototype of emotion embedding. We further learn the actual emotion encoder from an emotion-labelled database and study the use of relative attributes to represent fine-grained emotion intensity. To ensure emotional intelligibility, we incorporateemotion classification lossandemotion embedding similarity lossinto the training of the EVC network. As desired, the proposed network controls the fine-grained emotion intensity in the output speech. Through both objective and subjective evaluations, we validate the effectiveness of the proposed network for emotional expressiveness and emotion intensity control. Kun Zhou 0003, Berrak Sisman, Rajib Rana, Björn W. Schuller, Haizhou Li 0001 |
IEEE Trans. Affect. Comput. | 3 |
| 2023 | Speech Synthesis With Mixed EmotionsabstractEmotional speech synthesis aims to synthesize human voices with various emotional effects. The current studies are mostly focused on imitating an averaged style belonging to a specific emotion type. In this paper, we seek to generate speech with a mixture of emotions at run-time. We propose a novel formulation that measures the relative difference between the speech samples of different emotions. We then incorporate our formulation into a sequence-to-sequence emotional text-to-speech framework. During the training, the framework does not only explicitly characterize emotion styles but also explores the ordinal nature of emotions by quantifying the differences with other emotions. At run-time, we control the model to produce the desired emotion mixture by manually defining an emotion attribute vector. The objective and subjective evaluations have validated the effectiveness of the proposed framework. To our best knowledge, this research is the first study on modelling, synthesizing, and evaluating mixed emotions in speech. Kun Zhou 0003, Berrak Sisman, Rajib Rana, Björn W. Schuller, Haizhou Li 0001 |
IEEE Trans. Affect. Comput. | 3 |
| 2022 | Multi-Task Semi-Supervised Adversarial Autoencoding for Speech Emotion RecognitionabstractInspite the emerging importance of Speech Emotion Recognition (SER), the state-of-the-art accuracy is quite low and needs improvement to make commercial applications of SER viable. A key underlying reason for the low accuracy is the scarcity of emotion datasets, which is a challenge for developing any robust machine learning model in general. In this article, we propose a solution to this problem: a multi-task learning framework that uses auxiliary tasks for which data is abundantly available. We show that utilisation of this additional data can improve the primary task of SER for which only limited labelled data is available. In particular, we use gender identifications and speaker recognition as auxiliary tasks, which allow the use of very large datasets, e. g., speaker classification datasets. To maximise the benefit of multi-task learning, we further use an adversarial autoencoder (AAE) within our framework, which has a strong capability to learn powerful and discriminative features. Furthermore, the unsupervised AAE in combination with the supervised classification networks enables semi-supervised learning which incorporates a discriminative component in the AAE unsupervised training pipeline. This semi-supervised learning essentially helps to improve generalisation of our framework and thus leads to improvements in SER performance. The proposed model is rigorously evaluated for categorical and dimensional emotion, and cross-corpus scenarios. Experimental results demonstrate that the proposed model achieves state-of-the-art performance on two publicly available datasets. Siddique Latif, Rajib Rana, Sara Khalifa, Raja Jurdak, Julien Epps, Björn W. Schuller |
IEEE Trans. Affect. Comput. | 2 |
| 2021 | Guided Generative Adversarial Neural Network for Representation Learning and Audio Generation Using Fewer Labelled Audio DataabstractThe Generation power of Generative Adversarial Neural Networks (GANs) has shown great promise to learn representations from unlabelled data while guided by a small amount of labelled data. We aim to utilise the generation power of GANs to learn Audio Representations. Most existing studies are, however, focused on images. Some studies use GANs for speech generation, but they are conditioned on text or acoustic features, limiting their use for other audio, such as instruments, and even for speech where transcripts are limited. This paper proposes a novel GAN-based model that we named Guided Generative Adversarial Neural Network (GGAN), which can learn powerful representations and generate good-quality samples using a small amount of labelled data as guidance. Experimental results based on a speech [Speech Command Dataset (S09)] and a non-speech [Musical Instrument Sound dataset (Nsyth)] dataset demonstrate that using only 5% of labelled data as guidance, GGAN learns significantly better representations than the state-of-the-art models. Kazi Nazmul Haque, Rajib Rana, Jiajun Liu 0013, John H. L. Hansen, Nicholas Cummins, Carlos Busso, Björn W. Schuller |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2021 | Towards a Compressive-Sensing-Based Lightweight Encryption Scheme for the Internet of ThingsabstractInternet of Things (IoT) is flourishing and has penetrated deeply into people's daily life. With the seamless connection to the physical world, IoT provides tremendous opportunities to a wide range of applications. However, potential risks exist when the IoT system collects sensor data and uploads it to the Cloud. The leakage of private data can be severe with curious database administrator or malicious hackers who compromise the Cloud. In this work, we propose Kryptein, a compressive-sensing-based lightweight encryption scheme for Cloud-enabled IoT systems to secure the interaction between the IoT devices and the Cloud. Kryptein supports random compressed encryption, statistical computation over cipher, and accurate raw data decryption. According to our evaluation based on two real datasets, Kryptein provides strong protection to the data. It is 250 times faster than other state-of-the-art systems and incurs 120 times less energy consumption. The performance of Kryptein is also measured on off-the-shelf IoT devices, and the result shows Kryptein can run efficiently on IoT devices. After comparing with other state-of-the-art lightweight ciphers on IoT (Simon and Speck), IoT system with Kryptein is expected to have a much more longevity with about 35 percent extended lifetime. Further, experiments illustrated IoT data variance will not affect Kryptein's accuracy in a long term usage, and Krpytein is also able to support basic analytics tasks like machine learning (e.g., classification). Wanli Xue, Chengwen Luo 0001, Yiran Shen 0001, Rajib Rana, Guohao Lan, Sanjay K. Jha, Aruna Seneviratne, Wen Hu 0001 |
IEEE Trans. Mob. Comput. | 4 |
| 2020 | Augmenting Generative Adversarial Networks for Speech Emotion RecognitionabstractGenerative adversarial networks (GANs) have shown potential in learning emotional attributes and generating new data samples. However, their performance is usually hindered by the unavailability of larger speech emotion recognition (SER) data. In this work, we propose a framework that utilises the mixup data augmentation scheme to augment the GAN in feature learning and generation. To show the effectiveness of the proposed framework, we present results for SER on (i) synthetic feature vectors, (ii) augmentation of the training data with synthetic features, (iii) encoded features in compressed representation. Our results show that the proposed framework can effectively learn compressed emotional representations as well as it can generate synthetic samples that help improve performance in within-corpus and cross-corpus evaluation. Siddique Latif, Muhammad Asim 0005, Rajib Rana, Sara Khalifa, Raja Jurdak, Björn W. Schuller |
INTERSPEECH | 3 |
| 2020 | Deep Architecture Enhancing Robustness to Noise, Adversarial Attacks, and Cross-Corpus Setting for Speech Emotion RecognitionabstractSpeech emotion recognition systems (SER) can achieve high accuracy when the training and test data are identically distributed, but this assumption is frequently violated in practice and the performance of SER systems plummet against unforeseen data shifts. The design of robust models for accurate SER is challenging, which limits its use in practical applications. In this paper we propose a deeper neural network architecture wherein we fuse DenseNet, LSTM and Highway Network to learn powerful discriminative features which are robust to noise. We also propose data augmentation with our network architecture to further improve the robustness. We comprehensively evaluate the architecture coupled with data augmentation against (1) noise, (2) adversarial attacks and (3) cross-corpus settings. Our evaluations on the widely used IEMOCAP and MSP-IMPROV datasets show promising results when compared with existing studies and state-of-the-art models. Siddique Latif, Rajib Rana, Sara Khalifa, Raja Jurdak, Björn W. Schuller |
INTERSPEECH | 2 |
| 2020 | Poster Abstract: Federated Learning for Speech Emotion Recognition ApplicationsabstractPrivacy concerns are considered one of the major challenges in the applications of speech emotion recognition (SER) as it involves the complete sharing of speech data, which can bring threatening consequences to people’s lives. Federated learning is an effective technique to avoid privacy infringement by involving multiple participants to collaboratively learn a shared model without revealing their local data. In this work, we evaluated federated learning for SER using a publicly available dataset. Our preliminary results show that speech emotion recognition can benefit from federated learning by not exporting sensitive user data to central servers, while achieving promising results compared to the state-of-the-art. Siddique Latif, Sara Khalifa, Rajib Rana, Raja Jurdak |
IPSN | 3 |
| 2019 | Direct Modelling of Speech Emotion from Raw SpeechabstractSpeech emotion recognition is a challenging task and heavily depends on hand-engineered acoustic features, which are typically crafted to echo human perception of speech signals. However, a filter bank that is designed from perceptual evidence is not always guaranteed to be the best in a statistical modelling framework where the end goal is for example emotion classification. This has fuelled the emerging trend of learning representations from raw speech especially using deep learning neural networks. In particular, a combination of Convolution Neural Networks (CNNs) and Long Short Term Memory (LSTM) have gained great traction for the intrinsic property of LSTM in learning contextual information crucial for emotion recognition; and CNNs been used for its ability to overcome the scalability problem of regular neural networks. In this paper, we show that there are still opportunities to improve the performance of emotion recognition from the raw speech by exploiting the properties of CNN in modelling contextual information. We propose the use of parallel convolutional layers to harness multiple temporal resolutions in the feature extraction block that is jointly trained with the LSTM based classification network for the emotion recognition task. Our results suggest that the proposed model can reach the performance of CNN trained with hand-engineered features from both IEMOCAP and MSP-IMPROV datasets. Siddique Latif, Rajib Rana, Sara Khalifa, Raja Jurdak, Julien Epps |
INTERSPEECH | 2 |
| 2018 | Variational Autoencoders for Learning Latent Representations of Speech Emotion: A Preliminary StudyabstractLearning the latent representation of data in unsupervised fashion is a very interesting process that provides relevant features for enhancing the performance of a classifier. For speech emotion recognition tasks, generating effective features is crucial. Currently, handcrafted features are mostly used for speech emotion recognition, however, features learned automatically using deep learning have shown strong success in many problems, especially in image processing. In particular, deep generative models such as Variational Autoencoders (VAEs) have gained enormous success for generating features for natural images. Inspired by this, we propose VAEs for deriving the latent representation of speech signals and use this representation to classify emotions. To the best of our knowledge, we are the first to propose VAEs for speech emotion classification. Evaluations on the IEMOCAP dataset demonstrate that features learned by VAEs can produce state-of-the-art results for speech emotion classification. Siddique Latif, Rajib Rana, Junaid Qadir 0001, Julien Epps |
INTERSPEECH | 2 |
| 2018 | Transfer Learning for Improving Speech Emotion Classification AccuracyabstractThe majority of existing speech emotion recognition research focuses on automatic emotion detection using training and testing data from same corpus collected under the same conditions.The performance of such systems has been shown to drop significantly in cross-corpus and cross-language scenarios.To address the problem, this paper exploits a transfer learning technique to improve the performance of speech emotion recognition systems that is novel in cross-language and cross-corpus scenarios.Evaluations on five different corpora in three different languages show that Deep Belief Networks (DBNs) offer better accuracy than previous approaches on cross-corpus emotion recognition, relative to a Sparse Autoencoder and SVM baseline system.Results also suggest that using a large number of languages for training and using a small fraction of the target data in training can significantly boost accuracy compared with baseline also for the corpus with limited training examples. Siddique Latif, Rajib Rana, Shahzad Younis, Junaid Qadir 0001, Julien Epps |
INTERSPEECH | 2 |
| 2018 | A Novel Framework for Distress Detection through an Automated Speech Processing SystemabstractBased on our ongoing work, this work in progress project aims to develop an automated system to detect distress in people to enable early referral for interventions to target anxiety and depression, to mitigate suicidal ideation and to improve adherence to treatment. The project will utilize either use existing voice data to assess people into various scales of distress, or will collect voice data as per existing standards of distress measurement, to develop basic computing algorithms required to detect various attributes associated with distress, detected through a person's voice in a telephone call to a helpline. This will be then matched with the already available psychological assessment instruments such as the Distress Thermometer for these persons. In order to trigger interventions, organizational contexts are essential as interventions rely on the type of distress. Therefore, the model will be tested on various organizational settings such as the Police, Emergency and Health along with the Distress detection instruments normally used in a psychological assessment for accuracy and validation. The outcome of the project will culminate in a fully automated integrated system, and will save significant resources to organizations. The translation of the project will be realized in step-change improvements to quality of life within the gamut of public policy. Rajib Rana, Raj Gururajan, Geraldine Mackenzie, Jeff Dunn, Anthony Gray, Xujuan Zhou, Prabal Datta Barua, Julien Epps, Gerald Humphris |
WI | 1 |
| 2017 | Kryptein: a compressive-sensing-based encryption scheme for the internet of thingsabstractInternet of Things (IoT) is flourishing and has penetrated deeply into people's daily life. With the seamless connection to the physical world, IoT provides tremendous opportunities to a wide range of applications. However, potential risks exist when the IoT system collects sensor data and uploads it to the cloud. The leakage of private data can be severe with curious database administrator or malicious hackers who compromise the cloud. In this work, we propose Kryptein, a compressive-sensing-based encryption scheme for cloud-enabled IoT systems to secure the interaction between the IoT devices and the cloud. Kryptein supports random compressed encryption, statistical decryption, and accurate raw data decryption. According to our evaluation based on two real datasets, Kryptein provides strong protection to the data. It is 250 times faster than other state-of-the-art systems and incurs 120 times less energy consumption. The performance of Kryptein is also measured on off-the-shelf IoT devices, and the result shows Kryptein can run efficiently on IoT devices. Wanli Xue, Chengwen Luo 0001, Guohao Lan, Rajib Rana, Wen Hu 0001, Aruna Seneviratne |
IPSN | 4 |
| 2016 | A Model for Automated Affect Recognition on Smartphone-Cloud Architecture
Rajib Rana, Frank Whittaker, Jeffrey Soar |
ICOST | 2 |
| 2016 | CScrypt: A Compressive-Sensing-Based Encryption Engine for the Internet of Things: Demo AbstractabstractInternet of Things (IoT) have been connecting the physical world seamlessly and provides tremendous opportunities to a wide range of applications. However, potential risks exist when IoT system collects local sensor data and uploads to the Cloud. The private data leakage can be severe with curious database administrator or malicious hackers who compromise the Cloud. In this demo, we solve this problem of guaranteeing the user data privacy and security using compressive sensing based cryptographic method. We present CScrypt, a compressive-sensing-based encryption engine for the Cloud-enabled IoT systems to secure the interaction between the IoT devices and the Cloud. Our system exploits the fact that each individual's biometric data can be trained to a unique dictionary which can be used as an encryption key meanwhile to compress the original data. We will demonstrate a functioning prototype of our system using live data stream when attending the conference. Wanli Xue, Chengwen Luo 0001, Rajib Rana, Wen Hu 0001, Aruna Seneviratne |
SenSys | 3 |
| 2015 | Ear-Phone: A context-aware noise mapping using smart phones
Rajib Rana, Chun Tung Chou, Nirupama Bulusu, Salil S. Kanhere, Wen Hu 0001 |
Pervasive Mob. Comput. | 1 |
| 2013 | Real-time classification via sparse representation in acoustic sensor networksabstractAcoustic Sensor Networks (ASNs) have a wide range of applications in natural and urban environment monitoring, as well as indoor activity monitoring. In-network classification is critically important in ASNs because wireless transmission costs several orders of magnitude more energy than computation. The main challenges of in-network classification in ASNs include effective feature selection, intensive computation requirement and high noise levels. To address these challenges, we propose a sparse representation based feature-less, low computational cost, and noise resilient framework for in-network classification in ASNs. The key component of Sparse Approximation based Classification (SAC), ℓ1 minimization, is a convex optimization problem, and is known to be computationally expensive. Furthermore, SAC algorithms assumes that the test samples are a linear combination of a few training samples in the training sets. For acoustic applications, this results in a very large training dictionary, making the computation infeasible to be performed on resource constrained ASN platforms. Therefore, we propose several techniques to reduce the size of the problem, so as to fit SAC for in-network classification in ASNs. Our extensive evaluation using two real-life datasets (consisting of calls from 14 frog species and 20 cricket species respectively) shows that the proposed SAC framework outperforms conventional approaches such as Support Vector Machines (SVMs) and k-Nearest Neighbor (kNN) in terms of classification accuracy and robustness. Moreover, our SAC approach can deal with multi-label classification which is common in ASNs. Finally, we explore the system design spaces and demonstrate the real-time feasibility of the proposed framework by the implementation and evaluation of an acoustic classification application on an embedded ASN testbed. Bo Wei 0003, Mingrui Yang, Yiran Shen 0001, Rajib Rana, Chun Tung Chou, Wen Hu 0001 |
SenSys | 4 |
| 2012 | Distributed sparse approximation for frog sound classificationabstractSparse approximation has now become a buzzword for classification in numerous research domains. We propose a distributed sparse approximation method based on l1 minimization for frog sound classification, which is tailored to the resource constrained wireless sensor networks. Our pilot study demonstrates that l1 minimization can run on wireless sensor nodes producing satisfactory classification accuracy. Bo Wei 0003, Mingrui Yang, Rajib Rana, Chun Tung Chou, Wen Hu 0001 |
IPSN | 3 |
| 2011 | An Adaptive Algorithm for Compressive Approximation of Trajectory (AACAT) for Delay Tolerant Networks
Rajib Rana, Wen Hu 0001, Tim Wark, Chun Tung Chou |
EWSN | 1 |
| 2011 | Sparse Temporal Representations for Facial Expression Recognition
Sien W. Chew, Rajib Rana, Patrick Lucey, Simon Lucey, Sridha Sridharan |
PSIVT (2) | 2 |
| 2010 | Energy-Aware Sparse Approximation Technique (EAST) for Rechargeable Wireless Sensor Networks
Rajib Rana, Wen Hu 0001, Chun Tung Chou |
EWSN | 1 |
| 2010 | Ear-phone: an end-to-end participatory urban noise mapping systemabstractA noise map facilitates monitoring of environmental noise pollution in urban areas. It can raise citizen awareness of noise pollution levels, and aid in the development of mitigation strategies to cope with the adverse effects. However, state-of-the-art techniques for rendering noise maps in urban areas are expensive and rarely updated (months or even years), as they rely on population and traffic models rather than on real data. Participatory urban sensing can be leveraged to create an open and inexpensive platform for rendering up-to-date noise maps. Rajib Rana, Chun Tung Chou, Salil S. Kanhere, Nirupama Bulusu, Wen Hu 0001 |
IPSN | 1 |
| 2009 | Energy efficient information collection in wireless sensor networks using adaptive compressive sensingabstractWe consider the problem of using wireless sensor networks (WSNs) to measure the temporal-spatial field of some scalar physical quantities. Our goal is to obtain a sufficiently accurate approximation of the temporal-spatial field with as little energy as possible. We propose an adaptive algorithm, based on the recently developed theory of adaptive compressive sensing, to collect information from WSNs in an energy efficient manner. The key idea of the algorithm is to perform ¿projections¿ iteratively to maximise the amount of information gain per energy expenditure. We prove that this maximisation problem is NP-hard and propose a number of heuristics to solve this problem. We evaluate the performance of our proposed algorithms using data from both simulation and an outdoor WSN testbed. The results show that our proposed algorithms are able to give a more accurate approximation of the temporal-spatial field for a given energy expenditure. Chun Tung Chou, Rajib Rana, Wen Hu 0001 |
LCN | 2 |
| 2009 | Ear-Phone assessment of noise pollution with mobile phonesabstractNoise map can provide useful information to control noise pollution. We propose a people-centric noise collection system called the Ear-Phone. Due to the voluntary participation of people, the number and location of samples cannot be guaranteed. We propose and study two methods, based on compressive sensing, to reconstruct the missing samples. Rajib Rana, Chun Tung Chou, Salil S. Kanhere, Nirupama Bulusu, Wen Hu 0001 |
SenSys | 1 |