Juan Ye

dblp:75/3533 · DBLP profile ↗
← Back
44ranked-venue papers
12as first author
22since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Human-computer interaction and ubiquitous computing · 21 · 9 first-author · 6 since 2021Artificial intelligence and machine learning · 9 · 1 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 1 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Security and privacy · 1 · 1 since 2021Theory of computation · 1
YearPublicationVenuePosition
2026 The People's Gaze: Co-Designing and Refining Gaze Gestures with Users and Experts
abstract
As eye-tracking becomes increasingly common in modern mobile devices, the potential for hands-free, gaze-based interaction grows, but current gesture sets are largely expert-designed and often misaligned with how users naturally move their eyes. To address this gap, we introduce a two-phase methodology for developing intuitive gaze gestures. First, four co-design workshops with 20 non-expert participants generated 102 initial concepts. Next, four gaze interaction experts reviewed and refined these into a set of 32 gestures. We found that non-experts, after a brief introduction, intuitively anchor gestures in familiar metaphors and develop a compositional grammar; i.e., activation (dwell) + action (gaze gesture or blink), to ensure intentionality and mitigate the classic Midas Touch problem. Experts prioritized gestures that are ergonomically sound, aligned with natural saccades, and reliably distinguishable. The resulting user-grounded, expert-validated gesture set, along with actionable design principles, provides a foundation for developing intuitive, hands-free interfaces for gaze-enabled devices.
Yaxiong Lei, Xinya Gong, Shijing He, Yafei Wang 0004, Mohamed Khamis, Juan Ye
CHI6
2026 Privacy Perspectives and Practices of Chinese Smart Home Product Teams
abstract
Previous research has explored the privacy needs and concerns of device owners, primary users, and different bystander groups with regard to smart home devices like security cameras, smart speakers, and hubs, but little is known about the privacy views and practices of smart home product teams, particularly those in non-Western contexts. This paper presents findings from 27 semi-structured interviews with Chinese smart home product team members, including product/project managers, software/hardware engineers, user experience (UX) designers, legal/privacy experts, and marketers/operation specialists. We examine their privacy perspectives, practices, and risk mitigation strategies. Our results show that participants emphasized compliance with Chinese data privacy laws, which typically prioritized national security over individual privacy rights. China-specific cultural, social, and legal factors also influenced participants' ethical considerations and attitudes toward balancing user privacy and security with convenience. Drawing on our findings, we propose a set of recommendations for smart home product teams, along with socio-technical and legal interventions to address smart home privacy issues-especially those belonging to at-risk groups-in Chinese multi-user smart homes.
Shijing He, Yaxiong Lei, Xiao Zhan, Chi Zhang 0084, Juan Ye, Ruba Abu-Salma, Jose M. Such
SP5
2026 Fundus image quality assessment in retinopathy of prematurity via multi-label graph evidential network
Donghan Wu, Wenyue Shen, Heng Li 0010, Huaying Hao, Juan Ye, Yitian Zhao
Medical Image Anal.6
2026 Cross-domain neural alignment (CDNA): A deep supervised domain adaptation in human activity recognition using device-free sensing
abstract
Abstract Wi-Fi sensor networks have grown rapidly due to their scalability and high data throughput, finding applications in tasks like tracking human motion in laboratory settings. These systems detect motion by analyzing fluctuations in radio signals caused by target movements, which generate identifiable activity patterns. However, their performance is influenced by factors such as environmental changes, unseen target subjects, multi-target tracking, data configurations, and the nature of target activities. These challenges lead to domain shifts between the training and testing phases, a common issue in real-world scenarios known as the domain-shifting problem in transfer learning. We propose a supervised domain alignment technique to address domain shifts in Wi-Fi sensor Channel State Information (CSI) datasets using minimal labeled target data. Our method outperforms state-of-the-art adversarial models trained on similar data, achieving superior cross-domain prediction accuracy. Evaluations on two public CSI datasets show consistent improvements, with an average Micro-F1 score of 90% for cross-user tasks and 67% for cross-user and cross-environment tasks using only 70 labeled target samples. This approach demonstrates its effectiveness in enhancing prediction accuracy under challenging domain shift scenarios.
Thomas W. Kelsey, Andrea Rosales Sanabria, Juan Ye
Multim. Tools Appl.4
2025 LPUWF-LDM: Enhanced latent diffusion model for precise late-phase UWF-FA generation on limited dataset
Zhaojie Fang, Guanyu Zhou, Ke Zhuang, Yifei Chen 0019, Ruiquan Ge, Changmiao Wang, Gangyong Jia, Qing Wu 0008, Juan Ye, Maimaiti Nuliqiman, Peifang Xu, Ahmed El-Azab
Expert Syst. Appl.10
2025 Continual learning in sensor-based human activity recognition with dynamic mixture of experts
abstract
Human activity recognition (HAR) is a key enabler for many applications in healthcare, factory automation, and smart home. It detects and predicts human behaviours or daily activities via a range of wearable sensors or ambient sensors embedded in an environment. As more and more HAR applications are deployed in the real-world environments, there is a pressing need for the ability of continually and incrementally learning new activities over time without retraining the HAR model. Recently, various continual learning techniques have been applied to HAR; however, most of them commit to a large architecture, which might not suit to devices that deploy HAR models. In addition, these techniques often require to deploy the same large architecture on the devices and cannot customise the architecture for different requirements. To tackle this challenge, we present a dynamic mixture-of-experts approach, which grows an expert for each new task and allows flexible composition of experts to suit individual needs of applications. We have empirically evaluated our technique on 4 third-party, publicly available datasets and compared with 11 state-of-the-art continual learning techniques. Our results demonstrate that our technique can achieve better or comparable performance but with much less parameter spaces and training time.
Fahrurrozi Rahman, Martin Schiemer, Andrea Rosales Sanabria, Juan Ye
Pervasive Mob. Comput.4
2025 Unlabeled data augmentation with diffusion model for semi-supervised object detection
Zhanyun Lu, Renshu Gu, Huimin Cheng, Peifang Xu, Yuichiro Kinoshita, Juan Ye, Gangyong Jia, Qing Wu 0008
Vis. Comput.7
2024 Diffusers Generated Unlabeled Images Improves Semi-supervised Object Detection
abstract
In the field of object detection, particularly in medical imaging, the scarcity of data often poses a significant challenge to model performance. To address this issue, this study proposes a semi-supervised learning approach based on a generative model. We begin by fine-tuning a pre-trained generative model using our dataset to better adapt the generative model to our specific data distribution. The fine-tuned generative model is then used to generate additional unlabeled data. These generated unlabeled data, combined with the original dataset, are employed in a semi-supervised training process. Experimental results demonstrate that our method significantly enhances the performance of the object detection model, especially in scenarios with limited labeled data, such as medical imaging. By incorporating the generated unlabeled training data into the semi-supervised framework, we observed a notable improvement in model accuracy. Specifically, our experiments showed an increase of up to $6.92 \%$ after adding the generated iamges. Moreover, it is foreseeable that incorporating a higher proportion of generated unlabeled data could lead to even more significant improvements in performance.
Zhanyun Lu, Renshu Gu, Huimin Cheng, Peifang Xu, Yuichiro Kinoshita, Juan Ye, Gangyong Jia, Qing Wu 0008
CW7
2024 PGKD-Net: Prior-guided and Knowledge Diffusive Network for Choroid Segmentation
abstract
The thickness of the choroid is considered to be an important indicator of clinical diagnosis. Therefore, accurate choroid segmentation in retinal OCT images is crucial for monitoring various ophthalmic diseases. However, this is still challenging due to the blurry boundaries and interference from other lesions. To address these issues, we propose a novel prior-guided and knowledge diffusive network (PGKD-Net) to fully utilize retinal structural information to highlight choroidal region features and boost segmentation performance. Specifically, it is composed of two parts: a Prior-mask Guided Network (PG-Net) for coarse segmentation and a Knowledge Diffusive Network (KD-Net) for fine segmentation. In addition, we design two novel feature enhancement modules, Multi-Scale Context Aggregation (MSCA) and Multi-Level Feature Fusion (MLFF). The MSCA module captures the long-distance dependencies between features from different receptive fields and improves the model's ability to learn global context. The MLFF module integrates the cascaded context knowledge learned from PG-Net to benefit fine-level segmentation. Comprehensive experiments are conducted to evaluate the performance of the proposed PGKD-Net. Experimental results show that our proposed method achieves superior segmentation accuracy over other state-of-the-art methods. Our code is made up publicly available at: https://github.com/yzh-hdu/choroid-segmentation.
Yaqi Wang 0002, Zehua Yang, Xindi Liu, Dechao Chen, Gangyong Jia, Juan Ye, Xingru Huang
Artif. Intell. Medicine11
2024 Mmy-net: a multimodal network exploiting image and patient metadata for simultaneous segmentation and diagnosis
Renshu Gu, Yueyu Zhang, Lisha Wang, Dechao Chen, Yaqi Wang 0002, Ruiquan Ge, Zicheng Jiao, Juan Ye, Gangyong Jia, Linyan Wang
Multim. Syst.8
2024 Sketch-Supervised Histopathology Tumour Segmentation: Dual CNN-Transformer With Global Normalised CAM
abstract
Deep learning methods are frequently used in segmenting histopathology images with high-quality annotations nowadays. Compared with well-annotated data, coarse, scribbling-like labelling is more cost-effective and easier to obtain in clinical practice. The coarse annotations provide limited supervision, so employing them directly for segmentation network training remains challenging. We present a sketch-supervised method, called DCTGN-CAM, based on a dual CNN-Transformer network and a modified global normalised class activation map. By modelling global and local tumour features simultaneously, the dual CNN-Transformer network produces accurate patch-based tumour classification probabilities by training only on lightly annotated data. With the global normalised class activation map, more descriptive gradient-based representations of the histopathology images can be obtained, and inference of tumour segmentation can be performed with high accuracy. Additionally, we collect a private skin cancer dataset named BSS, which contains fine and coarse annotations for three types of cancer. To facilitate reproducible performance comparison, experts are also invited to label coarse annotations on the public liver cancer dataset PAIP2019. On the BSS dataset, our DCTGN-CAM segmentation outperforms the state-of-the-art methods and achieves 76.68 % IOU and 86.69 % Dice scores on the sketch-based tumour segmentation task. On the PAIP2019 dataset, our method achieves a Dice gain of 8.37 % compared with U-Net as the baseline network.
Yilong Li 0002, Linyan Wang, Xingru Huang, Yaqi Wang 0002, Ruiquan Ge, Huiyu Zhou 0001, Juan Ye, Qianni Zhang
IEEE J. Biomed. Health Informatics8
2023 GOMPS: Global Attention-based Ophthalmic Image Measurement and Postoperative Appearance Prediction System
abstract
Accurate measurements of ophthalmic parameters and postoperative appearance prediction are essential for the diagnosis and treatment of many ophthalmic diseases. Nevertheless, it remains challenging due to (1) inconsistent ophthalmic image sampling standards, including ocular-camera distance, facial angle, and patient number, (2) complicated ocular morphology, such as subconjunctival hemorrhage, ocular movements, lighting effects, and morphological aging. It is difficult for a model to measure parameters and make predictions in variable sampling methods and morphology conditions. Therefore, the Global attention-based Ophthalmic Image Measurement and Postoperative Appearance Prediction System (GOMPS) is proposed, which quantifies ophthalmic image parameters to diagnose disease and simultaneously predict postoperative appearance of blepharoptosis. By perceiving the global structure of the ophthalmic image, GOMPS makes logical inference predictions of the sclera and cornea morphology, to overcome the above difficulties. Concretely, a global attention unit (GAU) and a novel global attention structure-aware network (GASA-Net) are designed to enhance GOMPS’s global structure awareness ability to perform logical reasoning. Extensive experimental results on our collected ophthalmic dataset for diagnosis & prediction (OD2P) demonstrate that GOMPS surpasses the state-of-the-art methods in segmentation accuracy and achieves the current optimal performance in measurement and postoperative prediction under many clinical scenes.
Xingru Huang, Lixia Lou, Ruilong Dan, Lingxiao Chen, Guodong Zeng, Gangyong Jia, Qun Jin, Juan Ye, Yaqi Wang 0002
Expert Syst. Appl.10
2023 DynamicRead: Exploring Robust Gaze Interaction Methods for Reading on Handheld Mobile Devices under Dynamic Conditions
abstract
Enabling gaze interaction in real-time on handheld mobile devices has attracted significant attention in recent years. An increasing number of research projects have focused on sophisticated appearance-based deep learning models to enhance the precision of gaze estimation on smartphones. This inspires important research questions, including how the gaze can be used in a real-time application, and what type of gaze interaction methods are preferable under dynamic conditions in terms of both user acceptance and delivering reliable performance. To address these questions, we design four types of gaze scrolling techniques: three explicit technique based on Gaze Gesture, Dwell time, and Pursuit; and one implicit technique based on reading speed to support touch-free, page-scrolling on a reading application. We conduct a 20-participant user study under both sitting and walking settings and our results reveal that Gaze Gesture and Dwell time-based interfaces are more robust while walking and Gaze Gesture has achieved consistently good scores on usability while not causing high cognitive workload.
Yaxiong Lei, Tyler Caslin, Alexander Wisowaty, Xu Zhu 0005, Mohamed Khamis, Juan Ye
Proc. ACM Hum. Comput. Interact.7
2023 Online continual learning for human activity recognition
abstract
Sensor-based human activity recognition (HAR), with the ability to recognise human activities from wearable or embedded sensors, has been playing an important role in many applications including personal health monitoring, smart home, and manufacturing. The real-world, long-term deployment of these HAR systems drives a critical research question: how to evolve the HAR model automatically over time to accommodate changes in an environment or activity patterns. This paper presents an online continual learning (OCL) scenario for HAR, where sensor data arrives in a streaming manner which contains unlabelled samples from already learnt activities or new activities. We propose a technique, OCL-HAR, making a real-time prediction on the streaming sensor data while at the same time discovering and learning new activities. We have empirically evaluated OCL-HAR on four third-party, publicly available HAR datasets. Our results have shown that this OCL scenario is challenging to state-of-the-art continual learning techniques that have significantly underperformed. Our technique OCL-HAR has consistently outperformed them in all experiment setups, leading up to 0.17 and 0.23 improvements in micro and macro F1 scores.
Martin Schiemer, Lei Fang 0001, Simon A. Dobson, Juan Ye
Pervasive Mob. Comput.4
2023 Investigating Multisensory Integration in Emotion Recognition Through Bio-Inspired Computational Models
abstract
Emotion understanding represents a core aspect of human communication. Our social behaviours are closely linked to expressing our emotions and understanding others’ emotional and mental states through social signals. The majority of the existing work proceeds by extracting meaningful features from each modality and applying fusion techniques either at a feature level or decision level. However, these techniques are incapable of translating the constant talk and feedback between different modalities. Such constant talk is particularly important in continuous emotion recognition, where one modality can predict, enhance and complement the other. This article proposes three multisensory integration models, based on different pathways of multisensory integration in the brain; that is, integration by convergence, early cross-modal enhancement, and integration through neural synchrony. The proposed models are designed and implemented using third-generation neural networks, Spiking Neural Networks (SNN). The models are evaluated using widely adopted, third-party datasets and compared to state-of-the-art multimodal fusion techniques, such as early, late and deep learning fusion. Evaluation results show that the three proposed models have achieved comparable results to the state-of-the-art supervised learning techniques. More importantly, this article demonstrates plausible ways to translate constant talk between modalities during the training phase, which also brings advantages in generalisation and robustness to noise.
Esma Mansouri-Benssassi, Juan Ye
IEEE Trans. Affect. Comput.2
2023 CDNet: Contrastive Disentangled Network for Fine-Grained Image Categorization of Ocular B-Scan Ultrasound
abstract
Precise and rapid categorization of images in the B-scan ultrasound modality is vital for diagnosing ocular diseases. Nevertheless, distinguishing various diseases in ultrasound still challenges experienced ophthalmologists. Thus a novel contrastive disentangled network (CDNet) is developed in this work, aiming to tackle the fine-grained image categorization (FGIC) challenges of ocular abnormalities in ultrasound images, including intraocular tumor (IOT), retinal detachment (RD), posterior scleral staphyloma (PSS), and vitreous hemorrhage (VH). Three essential components of CDNet are the weakly-supervised lesion localization module (WSLL), contrastive multi-zoom (CMZ) strategy, and hyperspherical contrastive disentangled loss (HCD-Loss), respectively. These components facilitate feature disentanglement for fine-grained recognition in both the input and output aspects. The proposed CDNet is validated on our ZJU Ocular Ultrasound Dataset (ZJUOUSD), consisting of 5213 samples. Furthermore, the generalization ability of CDNet is validated on two public and widely-used chest X-ray FGIC benchmarks. Quantitative and qualitative results demonstrate the efficacy of our proposed CDNet, which achieves state-of-the-art performance in the FGIC task.
Ruilong Dan, Gangyong Jia, Shuai Wang 0003, Ruiquan Ge, Guiping Qian, Qun Jin, Juan Ye, Yaqi Wang 0002
IEEE J. Biomed. Health Informatics10
2022 A multi-feature deep learning system to enhance glaucoma severity diagnosis with high accuracy and fast speed
abstract
Glaucoma is the leading cause of irreversible blindness, and the early detection and timely treatment are essential for glaucoma management. However, due to the interindividual variability in the characteristics of glaucoma onset, a single feature is not yet sufficient for monitoring glaucoma progression in isolation. There is an urgent need to develop more comprehensive diagnostic methods with higher accuracy. In this study, we proposed a multi- feature deep learning (MFDL) system based on intraocular pressure (IOP), color fundus photograph (CFP) and visual field (VF) to classify the glaucoma into four severity levels. We designed a three-phase framework for glaucoma severity diagnosis from coarse to fine, which contains screening, detection and classification. We trained it on 6,131 samples from 3,324 patients and tested it on independent 240 samples from 185 patients. Our results show that MFDL achieved a higher accuracy of 0.842 (95 % CI, 0.795-0.888) than the direct four classification deep learning (DFC-DL, accuracy of 0.513 [0.449-0.576]), CFP-based single-feature deep learning (CFP-DL, accuracy of 0.483 [0.420-0.547]) and VF-based single-feature deep learning (VF-DL, accuracy of 0.725 [0.668-0.782]). Its performance was statistically significantly superior to that of 8 juniors. It also outperformed 3 seniors and 1 expert, and was comparable with 2 glaucoma experts (0.842 vs 0.854, p = 0.663; 0.842 vs 0.858, p = 0.580). With the assistance of MFDL, junior ophthalmologists achieved statistically significantly higher accuracy performance, with the increased accuracy ranged from 7.50 % to 17.9 %, and that of seniors and experts were 6.30 % to 7.50 % and 5.40 % to 7.50 %. The mean diagnosis time per patient of MFDL was 5.96 s. The proposed model can potentially assist ophthalmologists in efficient and accurate glaucoma diagnosis that could aid the clinical management of glaucoma.
Jiazhu Zhu, Yameng Zheng, Zhijing Zhu, Juan Ye, Ke Si
J. Biomed. Informatics9
2022 VisuaLizations As Intermediate Representations (VLAIR): An approach for applying deep learning-based computer vision to non-image-based data
abstract
Deep learning algorithms increasingly support automated systems in areas such as human activity recognition and purchase recommendation. We identify a current trend in which data is transformed first into abstract visualizations and then processed by a computer vision deep learning pipeline. We call this VisuaLization As Intermediate Representation (VLAIR) and believe that it can be instrumental to support accurate recognition in a number of fields while also enhancing humans’ ability to interpret deep learning models for debugging purposes or for personal use. In this paper we describe the potential advantages of this approach and explore various visualization mappings and deep learning architectures. We evaluate several VLAIR alternatives for a specific problem (human activity recognition in an apartment) and show that VLAIR attains classification accuracy above classical machine learning algorithms and several other non-image-based deep learning algorithms with several data representations.
Ai Jiang, Miguel A. Nacenta, Juan Ye
Vis. Informatics3
2021 Continual learning in sensor-based human activity recognition: An empirical benchmark analysis
Saurav Jha, Martin Schiemer, Franco Zambonelli, Juan Ye
Inf. Sci.4
2021 ContrasGAN: Unsupervised domain adaptation in Human Activity Recognition via adversarial and contrastive learning
Andrea Rosales Sanabria, Franco Zambonelli, Simon A. Dobson, Juan Ye
Pervasive Mob. Comput.4
2021 Generalisation and robustness investigation for facial and speech emotion recognition using bio-inspired spiking neural networks
abstract
Abstract Emotion recognition through facial expression and non-verbal speech represents an important area in affective computing. They have been extensively studied from classical feature extraction techniques to more recent deep learning approaches. However, most of these approaches face two major challenges: (1) robustness—in the face of degradation such as noise, can a model still make correct predictions? and (2) cross-dataset generalisation—when a model is trained on one dataset, can it be used to make inference on another dataset?. To directly address these challenges, we first propose the application of a spiking neural network (SNN) in predicting emotional states based on facial expression and speech data, then investigate, and compare their accuracy when facing data degradation or unseen new input. We evaluate our approach on third-party, publicly available datasets and compare to the state-of-the-art techniques. Our approach demonstrates robustness to noise, where it achieves an accuracy of 56.2% for facial expression recognition (FER) compared to 22.64% and 14.10% for CNN and SVM, respectively, when input images are degraded with the noise intensity of 0.5, and the highest accuracy of 74.3% for speech emotion recognition (SER) compared to 21.95% of CNN and 14.75% for SVM when audio white noise is applied. For generalisation, our approach achieves consistently high accuracy of 89% for FER and 70% for SER in cross-dataset evaluation and suggests that it can learn more effective feature representations, which lead to good generalisation of facial features and vocal characteristics across subjects.
Esma Mansouri-Benssassi, Juan Ye
Soft Comput.2
2021 Continual Activity Recognition with Generative Adversarial Networks
abstract
Continual learning is an emerging research challenge in human activity recognition (HAR). As an increasing number of HAR applications are deployed in real-world environments, it is important and essential to extend the activity model to adapt to the change in people’s activity routine. Otherwise, HAR applications can become obsolete and fail to deliver activity-aware services. The existing research in HAR has focused on detecting abnormal sensor events or new activities, however, extending the activity model is currently under-explored. To directly tackle this challenge, we build on the recent advance in the area of lifelong machine learning and design a continual activity recognition system, called HAR-GAN , to grow the activity model over time. HAR-GAN does not require a prior knowledge on what new activity classes might be and it does not require to store historical data by leveraging the use of Generative Adversarial Networks (GAN) to generate sensor data on the previously learned activities. We have evaluated HAR-GAN on four third-party, public datasets collected on binary sensors and accelerometers. Our extensive empirical results demonstrate the effectiveness of HAR-GAN in continual activity recognition and shed insight on the future challenges.
Juan Ye, Pakawat Nakwijit, Martin Schiemer, Saurav Jha, Franco Zambonelli
ACM Trans. Internet Things1
2020 Synch-Graph: Multisensory Emotion Recognition Through Neural Synchrony via Graph Convolutional Networks
abstract
Human emotions are essentially multisensory, where emotional states are conveyed through multiple modalities such as facial expression, body language, and non-verbal and verbal signals. Therefore having multimodal or multisensory learning is crucial for recognising emotions and interpreting social signals. Existing multisensory emotion recognition approaches focus on extracting features on each modality, while ignoring the importance of constant interaction and co-learning between modalities. In this paper, we present a novel bio-inspired approach based on neural synchrony in audio-visual multisensory integration in the brain, named Synch-Graph. We model multisensory interaction using spiking neural networks (SNN) and explore the use of Graph Convolutional Networks (GCN) to represent and learn neural synchrony patterns. We hypothesise that modelling interactions between modalities will improve the accuracy of emotion recognition. We have evaluated Synch-Graph on two state-of-the-art datasets and achieved an overall accuracy of 98.3% and 96.82%, which are significantly higher than the existing techniques.
Esma Mansouri-Benssassi, Juan Ye
AAAI2
2020 Unsupervised domain adaptation for activity recognition across heterogeneous datasets
Andrea Rosales Sanabria, Juan Ye
Pervasive Mob. Comput.2
2020 XLearn: Learning Activity Labels across Heterogeneous Datasets
abstract
Sensor-driven systems often need to map sensed data into meaningfully labelled activities to classify the phenomena being observed. A motivating and challenging example comes from human activity recognition in which smart home and other datasets are used to classify human activities to support applications such as ambient assisted living, health monitoring, and behavioural intervention. Building a robust and meaningful classifier needs annotated ground truth, labelled with what activities are actually being observed—and acquiring high-quality, detailed, continuous annotations remains a challenging, time-consuming, and error-prone task, despite considerable attention in the literature. In this article, we use knowledge-driven ensemble learning to develop a technique that can combine classifiers built from individually labelled datasets, even when the labels are sparse and heterogeneous. The technique both relieves individual users of the burden of annotation and allows activities to be learned individually and then transferred to a general classifier. We evaluate our approach using four third-party, real-world smart home datasets and show that it enhances activity recognition accuracies even when given only a very small amount of training data.
Juan Ye, Simon A. Dobson, Franco Zambonelli
ACM Trans. Intell. Syst. Technol.1
2020 Discovery and Recognition of Emerging Human Activities Using a Hierarchical Mixture of Directional Statistical Models
abstract
Human activity recognition plays a significant role in enabling pervasive applications as it abstracts low-level noisy sensor data into high-level human activities, which applications can respond to. With more and more activity-aware applications deployed in real-world environments, a research challenge emerges-discovering and learning new activities that have not been pre-defined or observed in the training phase. This paper tackles this challenge by proposing a hierarchical mixture of directional statistical models. The model supports incrementally, continuously updating the activity model over time with the reduced annotation effort and without the need for storing historical sensor data. We have validated this solution on four publicly available, third-party smart home datasets, and have demonstrated up to 91.5 percent accuracies of detecting and recognising new activities.
Lei Fang 0001, Juan Ye, Simon A. Dobson
IEEE Trans. Knowl. Data Eng.2
2019 Sensor-Based Human Activity Mining Using Dirichlet Process Mixtures of Directional Statistical Models
abstract
We have witnessed an increasing number of activity-aware applications being deployed in real-world environments, including smart home and mobile healthcare. The key enabler to these applications is sensor-based human activity recognition; that is, recognising and analysing human daily activities from wearable and ambient sensors. With the power of machine learning we can recognise complex correlations between various types of sensor data and the activities being observed. However the challenges still remain: (1) they often rely on a large amount of labelled training data to build the model, and (2) they cannot dynamically adapt the model with emerging or changing activity patterns over time. To directly address these challenges, we propose a Bayesian nonparametric model, i.e. Dirichlet process mixture of conditionally independent von Mises Fisher models, to enable both unsupervised and semi-supervised dynamic learning of human activities. The Bayesian nonparametric model can dynamically adapt itself to the evolving activity patterns without human intervention and the learning results can be used to alleviate the annotation effort. We evaluate our approach against real-world, third-party smart home datasets, and demonstrate significant improvements over the state-of-the-art techniques in both unsupervised and supervised settings.
Lei Fang 0001, Juan Ye, Simon A. Dobson
DSAA2
2019 Speech Emotion Recognition With Early Visual Cross-modal Enhancement Using Spiking Neural Networks
abstract
Speech emotion recognition (SER) is an important part of affective computing and signal processing research areas. A number of approaches, especially deep learning techniques, have achieved promising results on SER. However, there are still challenges in translating temporal and dynamic changes in emotions through speech. Spiking Neural Networks (SNN) have demonstrated as a promising approach in machine learning and pattern recognition tasks such as handwriting and facial expression recognition. In this paper, we investigate the use of SNNs for SER tasks and more importantly we propose a new cross-modal enhancement approach. This method is inspired by the auditory information processing in the brain where auditory information is preceded, enhanced and predicted by a visual processing in multisensory audio-visual processing. We have conducted experiments on two datasets to compare our approach with the state-of-the-art SER techniques in both uni-modal and multi-modal aspects. The results have demonstrated that SNNs can be an ideal candidate for modeling temporal relationships in speech features and our cross-modal approach can significantly improve the accuracy of SER.
Esma Mansouri-Benssassi, Juan Ye
IJCNN2
2019 Robust Audio Sensing with Multi-Sound Classification
abstract
Audio data is a highly rich form of information, often containing patterns with unique acoustic signatures. In pervasive sensing environments, because of the empowered smart devices, we have witnessed an increasing research interest in sound sensing to detect ambient environment, recognise users' daily activities, and infer their health conditions. However, the main challenge is that the real-world environment often contains multiple sound sources, which can significantly compromise the robustness of the above environment, event, and activity detection applications. In this paper, we explore different approaches in multi-sound classification, and propose a stacked classifier based on the recent advance in deep learning. We evaluate our proposed approach in a comprehensive set of experiments on both sound effect and real-world datasets. The results have demonstrated that our approach can robustly identify each sound category among mixed acoustic signals, without the need of any a priori knowledge about the number and signature of sounds in the mixed signals.
Peter Haubrick, Juan Ye
PerCom2
2019 Representation Learning for Minority and Subtle Activities in a Smart Home Environment
abstract
Daily human activity recognition using sensor data can be a fundamental task for many real-world applications, such as home monitoring and assisted living. One of the challenges in human activity recognition is to distinguish activities that have infrequent occurrence and less distinctive patterns. We propose a dissimilarity representation-based hierarchical classifier to perform two-phase learning. In the first phase, the classifier learns general features to recognise majority classes, and the second phase is to collect minority and subtle classes to identify fine difference between them. We compare our approach with a collection of state-of-the-art classification techniques on a real-world third-party dataset that is collected in a two-user home setting. Our results demonstrate that our hierarchical classifier approach outperforms the existing techniques in distinguishing users in performing the same type of activities. The key novelty of our approach is the exploration of dissimilarity representations and hierarchical classifiers, which allows us to highlight the difference between activities with subtle difference, and thus allows the identification of well-discriminating features.
Andrea Rosales Sanabria, Thomas W. Kelsey, Juan Ye
PerCom3
2018 SLearn: Shared learning human activity labels across multiple datasets
abstract
The research of sensor-based human activity recognition has been attracting increasing attention over years as it is playing an important role in various human-beneficiary applications such as ambient assistive living, health monitoring, and behaviour changing. Nowadays, the advancement of sensing and communication technologies has led to the possibility of collecting a large amount of sensor data, however, to build a reliable computational model and accurately recognise human activities we still need the annotations on sensor data. Acquiring high-quality, detailed, continuous annotations is a challenging task. In this paper, we explore the solution space on sharing annotated activities across different datasets in order to enhance the recognition accuracies. We have designed and developed two approaches: sharing training data and sharing classifiers towards addressing this challenge. We have validated the approach on three datasets and demonstrated their effectiveness in recognising activities only with annotations from as little as 0.1% of each dataset.
Juan Ye
PerCom1
2016 The Effect of Privacy Concerns on Privacy Recommenders
abstract
Location-sharing services such as Facebook and Foursquare/Swarm have become increasingly popular, due to the ease at which users can share their locations, and participate in services, games and other applications that leverage these locations. But it is important for people who use these services to configure appropriate location-privacy preferences so that they can control to whom they want to share their location information. Manually configuring these preferences may be burdensome and confusing, and so location-privacy preference recommenders based on crowdsourcing preferences from other users have been proposed. Whether people will accept the recommended preferences acquired from other users, who they may not know or trust, has not, however, been investigated. In this paper, we present a user experiment (n=99) to explore what factors influence people's acceptance of location-privacy preference recommenders. We find that 44% of our participants have privacy concerns about such recommenders. These concerns are shown to have a negative effect (p<0.001) on their acceptance of the recommendations and their satisfaction about their choices. Furthermore, users' acceptance of recommenders varies according to both context and recommendations being made. Our findings are potentially useful to designers of location-sharing services and privacy recommenders.
Juan Ye, Tristan Henderson
IUI2
2016 Detecting abnormal events on binary sensors in smart home environments
Juan Ye, Graeme Stevenson, Simon A. Dobson
Pervasive Mob. Comput.1
2016 Human Visual System-Based Fundus Image Quality Assessment of Portable Fundus Camera Photographs
abstract
Telemedicine and the medical "big data" era in ophthalmology highlight the use of non-mydriatic ocular fundus photography, which has given rise to indispensable applications of portable fundus cameras. However, in the case of portable fundus photography, non-mydriatic image quality is more vulnerable to distortions, such as uneven illumination, color distortion, blur, and low contrast. Such distortions are called generic quality distortions. This paper proposes an algorithm capable of selecting images of fair generic quality that would be especially useful to assist inexperienced individuals in collecting meaningful and interpretable data with consistency. The algorithm is based on three characteristics of the human visual system--multi-channel sensation, just noticeable blur, and the contrast sensitivity function to detect illumination and color distortion, blur, and low contrast distortion, respectively. A total of 536 retinal images, 280 from proprietary databases and 256 from public databases, were graded independently by one senior and two junior ophthalmologists, such that three partial measures of quality and generic overall quality were classified into two categories. Binary classification was implemented by the support vector machine and the decision tree, and receiver operating characteristic (ROC) curves were obtained and plotted to analyze the performance of the proposed algorithm. The experimental results revealed that the generic overall quality classification achieved a sensitivity of 87.45% at a specificity of 91.66%, with an area under the ROC curve of 0.9452, indicating the value of applying the algorithm, which is based on the human vision system, to assess the image quality of non-mydriatic photography, especially for low-cost ophthalmological telemedicine applications.
Shaoze Wang, Haitong Lu, Chuming Cheng, Juan Ye, Dahong Qian
IEEE Trans. Medical Imaging5
2015 Fault detection for binary sensors in smart home environments
abstract
Experiments in assisted living confirm that such systems can provide context-aware services that enable occupants to remain active and independent. They also demonstrate that abnormal sensor events hamper the correct identification of critical (and potentially life-threatening) situations, and that existing learning, estimation, and time-based approaches are inaccurate and inflexible when applied to multiple people sharing a living space. We propose a technique that integrates the semantics of sensor readings with statistical outlier detection. We evaluate the technique against four real-world datasets that include multiple individuals, and show consistent rates of anomaly detection across different environments.
Juan Ye, Graeme Stevenson, Simon A. Dobson
PerCom1
2015 Semantic web technologies in pervasive computing: A survey and research roadmap
Juan Ye, Stamatia Dasiopoulou, Graeme Stevenson, Georgios Meditskos, Efstratios Kontopoulos, Ioannis Kompatsiaris, Simon A. Dobson
Pervasive Mob. Comput.1
2015 KCAR: A knowledge-driven approach for concurrent activity recognition
Juan Ye, Graeme Stevenson, Simon A. Dobson
Pervasive Mob. Comput.1
2015 Developing pervasive multi-agent systems with nature-inspired coordination
Franco Zambonelli, Andrea Omicini, Bernhard Anzengruber, Gabriella Castelli, Francesco L. De Angelis, Giovanna Di Marzo Serugendo, Simon A. Dobson, Jose Luis Fernandez-Marquez, Alois Ferscha, Marco Mamei, Stefano Mariani 0001, Ambra Molesini, Sara Montagna, Jussi Nieminen, Danilo Pianini, Matteo Risoldi, Alberto Rosi, Graeme Stevenson, Mirko Viroli, Juan Ye
Pervasive Mob. Comput.20
2014 Privacy-aware location privacy preference recommendations
abstract
Location-Based Services have become increasingly popular due to the prevalence of smart devices and location-sharing applications such as Facebook and Foursquare. The protection of people's sensitive location data in such applications is an important requirement. Conventional location privacy protec
Juan Ye, Tristan Henderson
MobiQuitous2
2014 USMART: An Unsupervised Semantic Mining Activity Recognition Technique
abstract
Recognising high-level human activities from low-level sensor data is a crucial driver for pervasive systems that wish to provide seamless and distraction-free support for users engaged in normal activities. Research in this area has grown alongside advances in sensing and communications, and experiments have yielded sensor traces coupled with ground truth annotations about the underlying environmental conditions and user actions. Traditional machine learning has had some success in recognising human activities; but the need for large volumes of annotated data and the danger of overfitting to specific conditions represent challenges in connection with the building of models applicable to a wide range of users, activities, and environments. We present USMART, a novel unsupervised technique that combines data- and knowledge-driven techniques. USMART uses a general ontology model to represent domain knowledge that can be reused across different environments and users, and we augment a range of learning techniques with ontological semantics to facilitate the unsupervised discovery of patterns in how each user performs daily activities. We evaluate our approach against four real-world third-party datasets featuring different user populations and sensor configurations, and we find that USMART achieves up to 97.5% accuracy in recognising daily activities.
Juan Ye, Graeme Stevenson, Simon A. Dobson
ACM Trans. Interact. Intell. Syst.1
2012 Situation identification techniques in pervasive computing: A review
Juan Ye, Simon A. Dobson, Susan McKeever
Pervasive Mob. Comput.1
2011 A top-level ontology for smart environments
Juan Ye, Graeme Stevenson, Simon A. Dobson
Pervasive Mob. Comput.1
2009 Using Situation Lattices in Sensor Analysis
abstract
Highly sensorised systems present two parallel challenges: how to design a sensor suite that can efficiently and cost-effectively support the needs of given services; and to extract the semantically relevant interpretations, or ldquosituationsrdquo, from the flood of context data collected by the sensors. We describe mathematical structures called situation lattices that can be used to address these two problems simultaneously, allowing designers to both design and refine situation identification whilst offering insights into the design of sensor suites. We validate the accuracy and efficiency of our technique against a third-party data set and demonstrate how it can be used to evaluate sensor suite designs.
Juan Ye, Lorcan Coyle, Simon A. Dobson, Paddy Nixon
PerCom1
2009 Human-Behaviour Study with Situation Lattices
abstract
Most research in the area of smart environments focuses on improving the accuracy with which human activities can be recognised. Relatively little research has been done into how designers can gain insights into the behaviours their systems are observing, and feed these insights back into improving systems design. We describe a mathematical structure, the situation lattice, and show how it can be used to discover knowledge about activities and the way in which they can be sensed. We show how this knowledge can be used to improve activity recognition, using the example of a real-world smart home data set.
Juan Ye, Simon A. Dobson
SMC1