EDBT 2026 Demo / reviewers in the wild / expert
Hong Jia
dblp:57/3548
· DBLP profile ↗
39ranked-venue papers
12as first author
29since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 16 · 7 first-author · 9 since 2021Computer networks · 13 · 2 first-author · 12 since 2021Databases, data management, data science and information retrieval · 7 · 2 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 3 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021Systems, architecture and hardware · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Spiking Graph Predictive Coding for Reliable OOD Generalization
Jing Ren 0001, Jiapeng Du, Bowen Li 0012, Ziqi Xu 0001, Xin Zheng 0008, Hong Jia, Suyu Ma, Xiwei Xu 0001, Feng Xia 0001 |
WWW | 6 |
| 2026 | NarrativeSense: Predicting Affective States in University Students through Smartphone Sensing and Contextual NarrativesabstractMental health challenges are increasingly prevalent among university students, yet often go undetected due to reliance on traditional assessments that are subjective, infrequent, and lack behavioral context. Digital phenotyping through passively collected smartphone data offers a scalable alternative, but existing approaches often fail to integrate predictive accuracy with narrative-based insights. To overcome these limitations, we present NarrativeSense, a novel framework that combines machine learning models with narrative-based descriptions of daily life events inferred from smartphone sensing data to predict weekly affective states. The system incorporates language model components to transform behavioral patterns into contextualized, human-readable narratives that ground affective predictions in everyday experiences. This narrative layer complements structured prediction by offering intuitive, user-centered insights. Applied to longitudinal data from 58 university students over 119 days, NarrativeSense outperforms baseline machine learning models, standalone LLMs, and ensemble methods, while providing richer insights. Our findings demonstrate the potential of narrative-enhanced digital phenotyping for scalable and explainable mental health monitoring in educational and clinical settings. Yan Li 0186, Yihao Ding, Hong Jia, Vassilis Kostakos, Simon D'Alfonso |
ACM Trans. Comput. Heal. | 4 |
| 2026 | From prediction to explanation: Using screen text to understand smartphone use and user behaviourabstractSmartphones are essential to daily life, and their rich data streams have been used to study how people use their phones, and more broadly human behaviour. While previous research has largely focused on app usage and keystroke dynamics to predict smartphone use, these analyses are typically limited to making predictions rather than providing explanations or reasoning for observed behaviours. In this exploratory study, we investigate the potential of leveraging screen text and large language models (LLMs) to uncover insights and reasoning about user behaviour. Using a dataset of over 100 million on-screen words collected from 21 participants over two weeks, we explore multiple ways to use screen text and LLMs for three tasks: predicting the next app a user will open, inferring what real-world activities they are engaged in, and understanding how they interact within apps. Orthogonally, we demonstrate the interpretive capabilities of LLMs, highlighting their potential to explain the reasoning behind observed user actions. Our findings suggest that screen text holds promise for providing deeper insights into both digital and real-world human behaviour. We discuss the broader implications of our findings, including enhancing user experience and enabling privacy-preserving, on-device analysis, while proposing future research directions in screen text analysis. Songyan Teng, Hong Jia, Simon D'Alfonso, Vassilis Kostakos |
Int. J. Hum. Comput. Stud. | 2 |
| 2026 | A cascade framework for on-device uncertainty-aware event detection on microcontrollersabstractPervasive sensing enables diverse wearable event detection (WED) applications, but deploying machine learning models on resource-constrained microcontrollers (MCUs) poses significant challenges, particularly in ensuring prediction reliability under data shifts or out-of-distribution (OOD) inputs. While Uncertainty quantification methods offer a way to assess this reliability, many are computationally prohibitive for MCUs, and detecting multiple events concurrently further exacerbates resource constraints. Addressing these combined challenges, this paper presents an uncertainty and resource-aware framework designed for reliable and efficient multi-event WED on MCUs, significantly extending our preliminary work. The proposed framework achieves this by integrating Evidential Deep Learning (EDL) for efficient, single-pass uncertainty estimation with a novel cascade learning architecture. This architecture promotes resource efficiency via: (i) intra-event sharing using uncertainty-aware early exits within a staged model (shallow, medium, deep), allowing simpler samples to terminate inference earlier; and (ii) inter-event sharing using a multi-head design where multiple event detectors share a common backbone, minimizing overhead. System efficiency is further enhanced through MCU-specific optimizations, including targeted architecture search, quantization, efficient uncertainty operator implementation using standard TensorFlow Lite Micro (TFLM) operations, and library footprint reduction. We conducted extensive experiments on four distinct wearable datasets (Oesense, KWS, ECG5000, and HHAR) and two MCU platforms (STM32F446ZE, STM32H747XI), comparing the proposed framework against strong baselines including Deep Ensembles and Vanilla EDL. Results demonstrate the proposed framework’s effectiveness, achieving competitive accuracy and uncertainty performance (e.g., up to 22% lower NLL than data augmentation) while drastically reducing resource consumption, offering up to 8.64 × faster inference, up to 8.57 × lower energy use, and 55% smaller memory footprint compared to ensemble methods. The proposed framework enables the deployment of reliable, uncertainty-aware multi-event detection on a wider range of low-power MCUs. Hong Jia, Young D. Kwon, Dong Ma 0001, Nhat Pham, Lorena Qendro, Tam Vu 0001, Cecilia Mascolo |
Pervasive Mob. Comput. | 1 |
| 2025 | Token-Level Logits Matter: A Closer Look at Speech Foundation Models for Ambiguous Emotion Recognition
Jule Valendo Halim, Hong Jia, Ting Dang |
INTERSPEECH | 3 |
| 2025 | LiteFat: Lightweight Spatio-Temporal Graph Learning for Real-Time Driver Fatigue DetectionabstractDetecting driver fatigue is critical for road safety, as drowsy driving remains a leading cause of traffic accidents. Many existing solutions rely on computationally demanding deep learning models, which result in high latency and are unsuitable for embedded robotic devices with limited resources (such as intelligent vehicles/cars) where rapid detection is necessary to prevent accidents. This paper introduces LiteFat, a lightweight spatio-temporal graph learning model designed to detect driver fatigue efficiently while maintaining high accuracy and low computational demands. LiteFat involves converting streaming video data into spatio-temporal graphs (STG) using facial landmark detection, which focuses on key motion patterns and reduces unnecessary data processing. LiteFat uses MobileNet to extract facial features and create a feature matrix for the STG. A lightweight spatio-temporal graph neural network is then employed to identify signs of fatigue with minimal processing and low latency. Experimental results on benchmark datasets show that LiteFat performs competitively while significantly reduced computational complexity and latency as compared to current state-of-the-art methods. This work advances the development of real-time, resource-efficient human fatigue detection systems that can be implemented upon embedded robotic devices. Jing Ren 0001, Suyu Ma, Hong Jia, Xiwei Xu 0001, Ivan Lee 0001, Haytham Fayek, Xiaodong Li 0001, Feng Xia 0001 |
IROS | 3 |
| 2025 | E-BATS: Efficient Backpropagation-Free Test-Time Adaptation for Speech Foundation ModelsabstractSpeech Foundation Models encounter significant performance degradation when deployed in real-world scenarios involving acoustic domain shifts, such as background noise and speaker accents. Test-time adaptation (TTA) has recently emerged as a viable strategy to address such domain shifts at inference time without requiring access to source data or labels. However, existing TTA approaches, particularly those relying on backpropagation, are memory-intensive, limiting their applicability in speech tasks and resource-constrained settings. Although backpropagation-free methods offer improved efficiency, existing ones exhibit poor accuracy. This is because they are predominantly developed for vision tasks, which fundamentally differ from speech task formulations, noise characteristics, and model architecture, posing unique transferability challenges.
In this paper, we introduce E-BAT, first Efficient BAckpropagation-free TTA framework designed explicitly for speech foundation models. E-BAT achieves a balance between adaptation effectiveness and memory efficiency through three key components: (i) lightweight prompt adaptation for a forward-pass-based feature alignment, (ii) a multi-scale loss to capture both global (utterance-level) and local distribution shifts (token-level) and (iii) a test-time exponential moving average mechanism for stable adaptation across utterances. Experiments conducted on four noisy speech datasets spanning sixteen acoustic conditions demonstrate consistent improvements, with 4.1\%--13.5% accuracy gains over backpropogation-free baselines and 2.0$\times$–6.4$\times$ GPU memory savings compared to backpropogation-based methods. By enabling scalable and robust adaptation under acoustic variability, this work paves the way for developing more efficient adaptation approaches for practical speech processing systems in real-world environments. Jiaheng Dong, Hong Jia, Soumyajit Chatterjee, Abhirup Ghosh, Ting Dang |
NeurIPS | 2 |
| 2025 | FlowerTune: A Cross-Domain Benchmark for Federated Fine-Tuning of Large Language ModelsabstractLarge Language Models (LLMs) have achieved state-of-the-art results across diverse domains, yet their development remains reliant on vast amounts of publicly available data, raising concerns about data scarcity and the lack of access to domain-specific, sensitive information. Federated Learning (FL) presents a compelling framework to address these challenges by enabling decentralized fine-tuning on pre-trained LLMs without sharing raw data. However, the compatibility and performance of pre-trained LLMs in FL settings remain largely under explored. We introduce the FlowerTune LLM Leaderboard, a first-of-its-kind benchmarking suite designed to evaluate federated fine-tuning of LLMs across four diverse domains: general NLP, finance, medical, and coding. Each domain includes federated instruction-tuning datasets and domain-specific evaluation metrics. Our results, obtained through a collaborative, open-source and community-driven approach, provide the first comprehensive comparison across 26 pre-trained LLMs with different aggregation and fine-tuning strategies under federated settings, offering actionable insights into model performance, resource constraints, and domain adaptation. This work lays the foundation for developing privacy-preserving, domain-specialized LLMs for real-world applications. Yan Gao 0016, Massimo Roberto Scamarcia, Javier Fernández-Marqués, Mohammad Naseri, Chong Shen Ng, Dimitris Stripelis, Zexi Li 0001, Tao Shen 0002, Jiamu Bai, Daoyuan Chen, Zikai Zhang 0003, Rui Hu 0005, Inseo Song, Kangyoon Lee, Hong Jia, Ting Dang, Zheyuan Liu 0002, Daniel J. Beutel, Lingjuan Lyu, Nicholas D. Lane |
NeurIPS | 15 |
| 2025 | LightLLM: A Versatile Large Language Model for Predictive Light SensingabstractWe propose LightLLM, a model that fine tunes pre-trained large language models (LLMs) for light-based sensing tasks. It integrates a sensor data encoder to extract key features, a contextual prompt to provide environmental information, and a fusion layer to combine these inputs into a unified representation. This combined input is then processed by the pre-trained LLM, which remains frozen while being fine-tuned through the addition of lightweight, trainable components, allowing the model to adapt to new tasks without altering its original parameters. This approach enables flexible adaptation of LLM to specialized light sensing tasks with minimal computational overhead and retraining effort. We have implemented LightLLM for three light sensing tasks: light-based localization, outdoor solar forecasting, and indoor solar estimation. Using real-world experimental datasets, we demonstrate that LightLLM significantly outperforms state-of-the-art methods, achieving 4.4x improvement in localization accuracy and 3.4x improvement in indoor solar estimation when tested in previously unseen environments. We further demonstrate that LightLLM outperforms ChatGPT-4 with direct prompting, highlighting the advantages of LightLLM's specialized architecture for sensor data fusion with textual prompts. Hong Jia, Mahbub Hassan, Lina Yao 0001, Branislav Kusy, Wen Hu 0001 |
SenSys | 2 |
| 2025 | Leafeon: Toward Accurate Sensing of Leaf Water Content for Protected Cropping With mmWave RadarabstractPlant sensing plays an important role in modern smart agriculture and the farming industry. Remote radio sensing allows for monitoring essential indicators of plant health, such as leaf water content (WC). While recent studies have shown the potential of using millimeter-wave (mmWave) radar for plant sensing, many overlook crucial factors, such as leaf structure and surface roughness, which can impact the accuracy of the measurements. In this article, we introduce Leafeon, which leverages mmWave radar to measure leaf WC noninvasively. Utilizing electronic beam steering, multiple leaf perspectives are sent to a custom deep neural network, which discerns unique reflection patterns from subtle antenna variations, ensuring accurate and robust leaf WC estimations. We implement a prototype of Leafeon using a Commercial Off-The-Shelf mmWave radar and evaluate its performance with a variety of different leaf types. Leafeon was trained in-lab using high-resolution destructive leaf measurements, achieving a mean absolute error (MAE) of leaf WC as low as 3.17% for the Avocado leaf, significantly outperforming the state-of-the-art approaches with an MAE reduction of up to 55.7%. Furthermore, we conducted experiments on live plants in both indoor and glasshouse experimental farm environments. Our results showed a strong correlation between predicted leaf WC levels and drought events. Mark Cardamis, Hong Jia, Wenyao Chen, Yihe Yan, Oula Ghannoum, Aaron J. Quigley, Chun Tung Chou, Wen Hu 0001 |
IEEE Internet Things J. | 2 |
| 2025 | Categorical Data Clustering via Value Order Estimated Distance Metric LearningabstractClustering is a popular machine learning technique for data mining that can process and analyze datasets to automatically reveal sample distribution patterns. Since the ubiquitous categorical data naturally lack a well-defined metric space such as the Euclidean distance space of numerical data, the distribution of categorical data is usually under-represented, and thus valuable information can be easily twisted in clustering. This paper, therefore, introduces a novel order distance metric learning approach to intuitively represent categorical attribute values by learning their optimal order relationship and quantifying their distance in a line similar to that of the numerical attributes. Since subjectively created qualitative categorical values involve ambiguity and fuzziness, the order distance metric is learned in the context of clustering. Accordingly, a new joint learning paradigm is developed to alternatively perform clustering and order distance metric learning with low time complexity and a guarantee of convergence. Due to the clustering-friendly order learning mechanism and the homogeneous ordinal nature of the order distance and Euclidean distance, the proposed method achieves superior clustering accuracy on categorical and mixed datasets. More importantly, the learned order distance metric greatly reduces the difficulty of understanding and managing the non-intuitive categorical data. Experiments with ablation studies, significance tests, case studies, etc., have validated the efficacy of the proposed method. The source code is available at https://github.com/csmjzhao/OCL_Source_Code. Yiqun Zhang 0006, Mingjie Zhao 0003, Hong Jia, Mengke Li 0001, Yang Lu 0009, Yiu-Ming Cheung |
Proc. ACM Manag. Data | 3 |
| 2024 | StatioCL: Contrastive Learning for Time Series via Non-Stationary and Temporal ContrastabstractContrastive learning (CL) has emerged as a promising approach for representation learning in time series data by embedding similar pairs closely while distancing dissimilar ones. However, existing CL methods often introduce false negative pairs (FNPs) by neglecting inherent characteristics and then randomly selecting distinct segments as dissimilar pairs, leading to erroneous representation learning, reduced model performance, and overall inefficiency. To address these issues, we systematically define and categorize FNPs in time series into semantic false negative pairs and temporal false negative pairs for the first time: the former arising from overlooking similarities in label categories, which correlates with similarities in non-stationarity and the latter from neglecting temporal proximity. Moreover, we introduce StatioCL, a novel CL framework that captures non-stationarity and temporal dependency to mitigate both FNPs and rectify the inaccuracies in learned representations. By interpreting and differentiating non-stationary states, which reflect the correlation between trends or temporal dynamics with underlying data patterns, StatioCL effectively captures the semantic characteristics and eliminates semantic FNPs. Simultaneously, StatioCL establishes fine-grained similarity levels based on temporal dependencies to capture varying temporal proximity between segments and to mitigate temporal FNPs. Evaluated on real-world benchmark time series classification datasets, StatioCL demonstrates a substantial improvement over state-of-the-art CL methods, achieving a 2.9% increase in Recall and a 19.2% reduction in FNPs. Most importantly, StatioCL also shows enhanced data efficiency and robustness against label scarcity. Yu Wu 0021, Ting Dang, Dimitris Spathis, Hong Jia, Cecilia Mascolo |
CIKM | 4 |
| 2024 | Robust Categorical Data Clustering Guided by Multi-Granular Competitive LearningabstractData set composed of categorical features is very common in big data analysis tasks. Since categorical features are usually with a limited number of qualitative possible values, the nested granular cluster effect is prevalent in the implicit discrete distance space of categorical data. That is, data objects frequently overlap in space or subspace to form small compact clusters, and similar small clusters often form larger clusters. However, the distance space cannot be well-defined like the Euclidean distance due to the qualitative categorical data values, which brings great challenges to the cluster analysis of categorical data. In view of this, we design a Multi-Granular Competitive Penalization Learning (MGCPL) algorithm to allow potential clusters to interactively tune themselves and converge in stages with different numbers of naturally compact clusters. To leverage MGCPL, we also propose a Cluster Aggregation strategy based on MGCPL Encoding (CAME) to first encode the data objects according to the learned multi-granular distributions, and then perform final clustering on the embeddings. It turns out that the proposed MGCPL-guided Categorical Data Clustering (MCDC) approach is competent in automatically exploring the nested distribution of multi-granular clusters and highly robust to categorical data sets from various domains. Benefiting from its linear time complexity, MCDC is scalable to large-scale data sets and promising in pre-partitioning data sets or compute nodes for boosting distributed computing. Extensive experiments with statistical evidence demonstrate its superiority compared to state-of-the-art counterparts on various real public data sets. Shenghong Cai, Yiqun Zhang 0006, Xiaopeng Luo, Yiu-Ming Cheung, Hong Jia, Peng Liu 0045 |
ICDCS | 5 |
| 2024 | LiDARSpectra: Synthetic Indoor Spectral Mapping with Low-cost LiDARsabstractWe introduce LiDARSpectra, a novel approach utilizing mobile-integrated commodity Light Detection and Ranging (LiDAR) signals for synthetic indoor light spectral mapping. Our method incorporates an innovative material estimation algorithm into the LiDAR signal processing pipeline, accurately simulating reflected wavelengths from indoor surfaces. Utilizing low-resolution LiDAR scans enriched with material information, it eliminates the need for deploying dedicated spectral sensors, greatly simplifying the spectral mapping process. We validate our synthetic spectral maps against real sensor data and demonstrate their utility in applications such as indoor localization and solar energy provisioning. This presents an efficient solution for indoor spectral mapping with wide-ranging potential across fields like lighting design, indoor planting, environmental monitoring, and location-based services. Hong Jia, Mahbub Hassan, Branislav Kusy, Wen Hu 0001 |
IPSN | 3 |
| 2024 | Efficient and Personalized Mobile Health Event Prediction via Small Language ModelsabstractHealthcare monitoring is crucial for early detection, timely intervention, and the ongoing management of health conditions, ultimately improving individuals' quality of life. Recent research shows that Large Language Models (LLMs) have demonstrated impressive performance in supporting healthcare tasks. However, existing LLM-based healthcare solutions typically rely on cloud-based systems, which raise privacy concerns and increase the risk of personal information leakage. As a result, there is growing interest in running these models locally on devices like mobile phones and wearables to protect users' privacy. Small Language Models (SLMs) are potential candidates to solve privacy and computational issues, as they are more efficient and better suited for local deployment. However, the performance of SLMs in healthcare domains has not yet been investigated. This paper examines the capability of SLMs to accurately analyze health data, such as steps, calories, sleep minutes, and other vital statistics, to assess an individual's health status. Our results show that, TinyLlama, which has 1.1 billion parameters, utilizes 4.31 GB memory, and has 0.48s latency, showing the best performance compared other four state-of-the-art (SOTA) SLMs on various healthcare applications. Our results indicate that SLMs could potentially be deployed on wearable or mobile devices for real-time health monitoring, providing a practical solution for efficient and privacy-preserving healthcare. Xin Wang 0215, Ting Dang, Vassilis Kostakos, Hong Jia |
MobiCom | 4 |
| 2024 | AutoJournaling: A Context-Aware Journaling System Leveraging MLLMs on Smartphone ScreenshotsabstractJournaling offers significant benefits, including fostering self-reflection, enhancing writing skills, and aiding in mood monitoring. However, many people abandon the practice because traditional journaling is time-consuming, and detailed life events may be overlooked if not recorded promptly. Given that smartphones are the most widely used devices for entertainment, work, and socialization, they present an ideal platform for innovative approaches to journaling. Despite their ubiquity, the potential of using digital phenotyping, a method of unobtrusively collecting data from digital devices to gain insights into psychological and behavioral patterns, for automated journal generation has been largely underexplored. In this study, we propose AutoJournaling, the first-of-its-kind system that automatically generates journals by collecting and analyzing screenshots from smartphones. This system captures life events and corresponding emotions, offering a novel approach to digital phenotyping. We evaluated AutoJournaling by collecting screenshots every 3 seconds from three students over five days, demonstrating its feasibility and accuracy. AutoJournaling is the first framework to utilize seamlessly collected screenshots for journal generation, providing new insights into psychological states through digital phenotyping. Shiquan Zhang, Hong Jia, Vassilis Kostakos, Simon D'Alfonso |
MobiCom | 4 |
| 2024 | TinyTTA: Efficient Test-time Adaptation via Early-exit Ensembles on Edge DevicesabstractThe increased adoption of Internet of Things (IoT) devices has led to the generation of large data streams with applications in healthcare, sustainability, and robotics. In some cases, deep neural networks have been deployed directly on these resource-constrained units to limit communication overhead, increase efficiency and privacy, and enable real-time applications. However, a common challenge in this setting is the continuous adaptation of models necessary to accommodate changing environments, i.e., data distribution shifts. Test-time adaptation (TTA) has emerged as one potential solution, but its validity has yet to be explored in resource-constrained hardware settings, such as those involving microcontroller units (MCUs). TTA on constrained devices generally suffers from i) memory overhead due to the full backpropagation of a large pre-trained network, ii) lack of support for normalization layers on MCUs, and iii) either memory exhaustion with large batch sizes required for updating or poor performance with small batch sizes. In this paper, we propose TinyTTA, to enable, for the first time, efficient TTA on constrained devices with limited memory. To address the limited memory constraints, we introduce a novel self-ensemble and batch-agnostic early-exit strategy for TTA, which enables continuous adaptation with small batch sizes for reduced memory usage, handles distribution shifts, and improves latency efficiency. Moreover, we develop the TinyTTA Engine, a first-of-its-kind MCU library that enables on-device TTA. We validate TinyTTA on a Raspberry Pi Zero 2W and an STM32H747 MCU. Experimental results demonstrate that TinyTTA improves TTA accuracy by up to 57.6\%, reduces memory usage by up to six times, and achieves faster and more energy-efficient TTA. Notably, TinyTTA is the only framework able to run TTA on MCU STM32H747 with a 512 KB memory constraint while maintaining high performance. Hong Jia, Young D. Kwon, Alessio Orsino, Ting Dang, Domenico Talia, Cecilia Mascolo |
NeurIPS | 1 |
| 2024 | UR2M: Uncertainty and Resource-Aware Event Detection on MicrocontrollersabstractTraditional machine learning techniques are prone to generating inaccurate predictions when confronted with shifts in the distribution of data between the training and testing phases. This vulnerability can lead to severe consequences, especially in applications such as mobile healthcare. Uncertainty estimation has the potential to mitigate this issue by assessing the reliability of a model's output. However, existing uncertainty estimation techniques often require substantial computational resources and memory, making them impractical for implementation on microcontrollers (MCUs). This limitation hinders the feasibility of many important on-device wearable event detection (WED) applications, such as heart attack detection. In this paper, we present UR2M, a novel Uncertainty and Resource-aware event detection framework for MCUs. Specifically, we (i) develop an uncertainty-aware WED based on evidential theory for accurate event detection and reliable uncertainty estimation; (ii) introduce a cascade ML framework to achieve efficient model inference via early exits, by sharing shallower model layers among different event models; (iii) optimize the deployment of the model and MCU library for system efficiency. We conducted extensive experiments and compared UR2M to traditional uncertainty baselines using three wearable datasets. Our results demonstrate that UR2M achieves up to 864% faster inference speed, 857% energy-saving for uncertainty estimation, 55% memory saving on two popular MCUs, and a 22% improvement in uncertainty quantification performance. UR2M can be deployed on a wide range of MCUs, significantly expanding real-time and reliable WED applications. Hong Jia, Young D. Kwon, Dong Ma 0001, Nhat Pham, Lorena Qendro, Tam Vu 0001, Cecilia Mascolo |
PerCom | 1 |
| 2023 | Subspace Clustering with Feature Grouping for Categorical Data
Hong Jia, Menghan Dong |
KSEM (1) | 1 |
| 2023 | Low Redundancy Learning for Unsupervised Multi-view Feature Selection
Hong Jia |
KSEM (1) | 1 |
| 2023 | Ubiquitous, Secure, and Efficient Mobile Sensing SystemsabstractThe rapid development of mobile sensors and machine learning techniques has enabled the utilization of diverse sensor data in many applications. However, the design and implementation of these algorithms and systems to achieve ubiquitous, secure, and efficient real-world solutions remains challenging. This paper discusses various effective solutions to overcome these challenges and underscores potential future directions. Hong Jia |
MobiSys | 1 |
| 2023 | LifeLearner: Hardware-Aware Meta Continual Learning System for Embedded Computing PlatformsabstractContinual Learning (CL) allows applications such as user personalization and household robots to learn on the fly and adapt to context. This is an important feature when context, actions, and users change. However, enabling CL on resource-constrained embedded systems is challenging due to the limited labeled data, memory, and computing capacity. Young D. Kwon, Jagmohan Chauhan, Hong Jia, Stylianos I. Venieris, Cecilia Mascolo |
SenSys | 3 |
| 2023 | Pistis: Replay Attack and Liveness Detection for Gait-Based User Authentication System on Wearable Devices Using VibrationabstractWearable devices-based biometrics has become mainstream in the biometric domain, especially in mobile computing, due to its convenience, flexibility, and potentially high user acceptance. Among various modalities, wearable devices-based gait recognition has been recognized as an effective user authentication method and employed in various applications, such as automated entry systems for home, school, work, vehicles, and automated ticket payment/validation for public transport. However, how secure wearable gait remains an open research question. In this study, we conduct a comprehensive security analysis of the wearable gait. Then, we demonstrate that gait itself is not robust against some attacking methods, such as spoofing or forgery. Therefore, we argue that an anti-spoofing mechanism is important for enhancing the security of wearable gait biometric systems. To this end, we proposed a novel authentication protocol called$Pistis$that embedded gait biometrics and a liveness detection mechanism that is aiming to detect various attacks of gait authentication systems. Our extensive experiments based on 50 subjects demonstrate that$Pistis$is effective in liveness detection and authentication performance enhancement, providing 100% accuracy for human and nonhuman detection, and 99.53% accuracy for user authentication. Pistis can be used as a liveness detection method for wearable devices-based biometrics, significantly for wearable gait. Hong Jia, Min Wang 0009, Yuezhong Wu, Wanli Xue, Chun Tung Chou, Jiankun Hu, Wen Hu 0001 |
IEEE Internet Things J. | 2 |
| 2023 | Penalized logistic regressions with technical indicators predict up and down trends
Huifeng Jiang, Hong Jia |
Soft Comput. | 3 |
| 2023 | Subject-adaptive Loose-fitting Smart Garment Platform for Human Activity RecognitionabstractThe ability to recognize and detect changes in human posture is important in a wide range of applications such as health care and human–computer interaction. Achieving this goal using loose-fit garments instrumented with sensors is particularly challenging, due to the complex interaction between garments and human body. Herein we present a method to detect and recognize human posture with casual loose-fitting smart garments integrated with highly sensitive, stretchable, optical transparent, and low-cost strain sensors. By attaching these sensors to an off-the-shelf casual jacket, we developed a smart loose-fitting sensing garment that enables posture recognition using a deep learning model, domain-adaptive Convolutional Neural Networks–Long Short-Term Memory (CNN-LSTM). This deep learning model overcame the noise and variation due to the complex interaction between loose-fitting garments and human body. Considering that users’ labeled data are usually not available in the training stage, an additional domain discriminator path on the conventional CNN-LSTM model has been introduced to further improve the adaptability. To evaluate the potential of this loose-fitting smart garment, three case studies were conducted under realistic conditions: recognitions of human activities, stationary postures with random hand movements and slouch. Our results demonstrate the potential of the proposed smart garment system for practical applications. Shuhua Peng, Yuezhong Wu, Jun Liu 0074, Hong Jia, Wen Hu 0001, Mahbub Hassan, Aruna Seneviratne, Chun Hui Wang |
ACM Trans. Sens. Networks | 5 |
| 2022 | Passive light spectral indoor localizationabstractWe propose a novel Visible Light Positioning (VLP) method, called Iris, that uses light spectral information (LSI) to localize humans completely passively in the sense that it neither requires the user to carry any device, nor does it require any modifications to existing lighting infrastructure. Iris localizes a user based on the interference they produce on the LSI recorded at an array of spectral sensors embedded in the environment. We design a deep neural network that can effectively learn location fingerprints directly from the sensor LSI data and predict locations accurately under varying lighting conditions. We prototype Iris using a commercial-off-the-shelf light spectral sensor, AS7265x, which can measure light intensity over 18 different wavelength channels. We benchmark Iris against the state-of-the-art passive VLPs that rely on conventional photo-sensors capable of measuring only a single light intensity value aggregated over the entire visible spectrum. Our evaluations over two typical indoor environments, a 25 m2 one-bedroom apartment and a 13m × 8m office space, demonstrate that Iris can significantly reduce both the localization errors and the number of required sensors, while increasing robustness against changes in environmental lighting. Hong Jia, Wen Hu 0001, Mahbub Hassan, Ashraf Uddin 0002, Branislav Kusy, Moustafa Youssef 0001 |
MobiCom | 3 |
| 2022 | PROS: an efficient pattern-driven compressive sensing framework for low-power biopotential-based wearables with on-chip intelligenceabstractWhile the global healthcare market of wearable devices has been growing significantly in recent years and is predicted to reach $60 billion by 2028, many important healthcare applications such as seizure monitoring, drowsiness detection, etc. have not been deployed due to the limited battery lifetime, slow response rate, and inadequate biosignal quality. Nhat Pham, Hong Jia, Tuan Dinh, Nam Bui, Young D. Kwon, Dong Ma 0001, Phuc Nguyen 0002, Cecilia Mascolo, Tam Vu 0001 |
MobiCom | 2 |
| 2022 | Indoor localization using light spectral informationabstractIn this paper, we investigate the impacts of location on the spectral distribution of received light, i.e., the intensity of light for different wavelengths, in indoor environments. Our findings show that, even when using the same light source, different locations exhibit slightly different spectral distribution due to reflections from their localised environment containing different materials or colours. Based on this observation, we present Spectral-Loc, a novel indoor localization method that employs light spectrum information to detect the device's position. Because spectrum sensors are increasingly being used in new products and applications, such as white balance in smartphone photography, Spectral-Loc can be quickly implemented without the need for extra hardware or infrastructure. We used a commercially available light spectrum sensor, the AS7265x, to prototype Spectral-Loc, which can measure light intensity over 18 different wavelength sub-bands. We benchmark the localization accuracy of Spectral-Loc against the conventional light intensity sensors that provide only a single intensity value. Our evaluations in two indoor areas, a meeting room and a large office, show that using light spectral information considerably decreases the localization error for different percentiles. Hong Jia, Wen Hu 0001, Mahbub Hassan, Ashraf Uddin 0002, Branislav Kusy, Moustafa Youssef 0001 |
MobiCom | 3 |
| 2021 | Condor: Mobile Golf Swing Tracking via Sensor Fusion using Conditional Generative Adversarial Networks
Hong Jia, Jun Liu 0074, Yuezhong Wu, Tomasz Bednarz, Lina Yao 0001, Wen Hu 0001 |
EWSN | 1 |
| 2019 | Mobile golf swing tracking using deep learning with data fusion: poster abstractabstractSwing tracking is one of the key information for many sports such as golf. One approach to track swing is to use IMU to measure linear acceleration then get position by two-time integration. However, the complex noise model of the IMU limit the accuracy of the tracking. Another approach is to use depth sensor to measure 3D location of a point of interest directly. Unfortunately, the depth sensor-based approach cannot accurately measure the trajectory of a swing when the sensor is occluded, which happens regularly. To overcome these limitations, we develop a novel solution to make use of these two sensor modalities (i.e., IMU and depth sensor) by a novel deep neural network to produce high precision swing trajectory tracking. The learned network automatically makes use of the IMU when the depth sensor is occluded, and relies on depth sensor when IMU signal is noisy. Our experiment shows that the proposed method outperforms state-of-the-art swing tracking method by 62% of error reduction. Hong Jia, Yuezhong Wu, Jun Liu 0074, Lina Yao 0001, Wen Hu 0001 |
SenSys | 1 |
| 2018 | Subspace Clustering of Categorical and Numerical Data With an Unknown Number of ClustersabstractIn clustering analysis, data attributes may have different contributions to the detection of various clusters. To solve this problem, the subspace clustering technique has been developed, which aims at grouping the data objects into clusters based on the subsets of attributes rather than the entire data space. However, the most existing subspace clustering methods are only applicable to either numerical or categorical data, but not both. This paper, therefore, studies the soft subspace clustering of data with both of the numerical and categorical attributes (also simply called mixed data for short). Specifically, an attribute-weighted clustering model based on the definition of object-cluster similarity is presented. Accordingly, a unified weighting scheme for the numerical and categorical attributes is proposed, which quantifies the attribute-to-cluster contribution by taking into account both of intercluster difference and intracluster similarity. Moreover, a rival penalized competitive learning mechanism is further introduced into the proposed soft subspace clustering algorithm so that the subspace cluster structure as well as the most appropriate number of clusters can be learned simultaneously in a single learning paradigm. In addition, an initialization-oriented method is also presented, which can effectively improve the stability and accuracy of -means-type clustering methods on numerical, categorical, and mixed data. The experimental results on different benchmark data sets show the efficacy of the proposed approach. Hong Jia, Yiu-Ming Cheung |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2016 | Superpixel-based coastline extraction in SAR images with speckle noise removalabstractCoastline extraction in Synthetic aperture radar (SAR) images is a fundamental and challenging task due to the speckle noise. In this paper, we propose a new method for automatic coastline extraction in SAR images. In our method, we combine K-means and speckle noise removal methods together to increase the dissimilarity between sea and land. To enhance the robustness to speckle noise, and preserve the targets boundaries, we treat superpixels as basic regions instead of pixels in traditional pixel-based methods. Finally, an adaptive threshold is applied to classify these regions into sea or land. Based on the classifications, a canny detector is employed to detect the coastline. We evaluate our proposed method on SAR images and the improved coastline extraction method superpixel-based is verified on remote sensing images with RGB channels. The experimental results demonstrate its superior performance on coastline extraction. Xiaofang Liu, Hong Jia, Liujuan Cao, Cheng Wang 0003, Jonathan Li 0001, Ming Cheng 0002 |
IGARSS | 2 |
| 2016 | A New Distance Metric for Unsupervised Learning of Categorical DataabstractDistance metric is the basis of many learning algorithms, and its effectiveness usually has a significant influence on the learning results. In general, measuring distance for numerical data is a tractable task, but it could be a nontrivial problem for categorical data sets. This paper, therefore, presents a new distance metric for categorical data based on the characteristics of categorical values. In particular, the distance between two values from one attribute measured by this metric is determined by both the frequency probabilities of these two values and the values of other attributes that have high interdependence with the calculated one. Dynamic attribute weight is further designed to adjust the contribution of each attribute-distance to the distance between the whole data objects. Promising experimental results on different real data sets have shown the effectiveness of the proposed distance metric. Hong Jia, Yiu-Ming Cheung, Jiming Liu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2014 | A new distance metric for unsupervised learning of categorical dataabstractDistance metric is the basis of many learning algorithms and its effectiveness usually has significant influence on the learning results. Generally, measuring distance for numerical data is a tractable task, but for categorical data sets, it could be a nontrivial problem. This paper therefore presents a new distance metric for categorical data based on the characteristics of categorical values. Specifically, the distance between two values from one attribute measured by this metric is determined by both of the frequency probabilities of these two values and the values of other attributes which have high interdependency with the calculated one. Promising experimental results on different real data sets have shown the effectiveness of proposed distance metric. Hong Jia, Yiu-Ming Cheung |
IJCNN | 1 |
| 2014 | Cooperative and penalized competitive learning with application to kernel-based clustering
Hong Jia, Yiu-Ming Cheung, Jiming Liu 0001 |
Pattern Recognit. | 1 |
| 2013 | A Unified Metric for Categorical and Numerical Attributes in Data Clustering
Yiu-Ming Cheung, Hong Jia |
PAKDD (2) | 2 |
| 2013 | Categorical-and-numerical-attribute data clustering based on a unified similarity metric without knowing cluster number
Yiu-Ming Cheung, Hong Jia |
Pattern Recognit. | 2 |
| 2012 | Unsupervised Feature Selection with Feature ClusteringabstractAs an effective technique for dimensionality reduction, feature selection has a broad application in different research areas. In this paper, we present a feature selection method based on a novel feature clustering procedure, which aims at partitioning the features into different clusters such that the features in the same cluster contain similar structural information of the given instances. Subsequently, since the obtained feature subset consists of features from variant clusters, the similarity between selected features will be low. This allows us to reserve the most data structural information with the minimum number of features. Experimental results on different benchmark data sets demonstrate the superiority of the proposed method. Yiu-Ming Cheung, Hong Jia |
Web Intelligence | 2 |
| 2010 | A Cooperative and Penalized Competitive Learning Approach to Gaussian Mixture Clustering
Yiu-Ming Cheung, Hong Jia |
ICANN (3) | 2 |