Hongkai Chen 0001

dblp:36/8493-1 · DBLP profile ↗
← Back
14ranked-venue papers
2as first author
12since 2021 · last 2026
0000-0001-7206-6584ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 10 · 9 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Theory of computation · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 A Large-Scale Multimodal Dataset and Benchmarks for Human Activity Scene Understanding and Reasoning
abstract
Multimodal human action recognition (HAR) utilizes complementary data for activity classification. Built on traditional HAR tasks, recent advances in Large Language Models (LLMs) enable detailed descriptions and causal reasoning of human actions, advancing new tasks of human action understanding (HAU) and human action reasoning (HARn). However, most LLMs, especially multimodal Large Vision-Language Models (LVLMs), struggle with modalities other than RGB images, like depth, IMU, ormmWave, due to a lack of large-scale datasets in these task domains. Existing HAR datasets provide only coarse-grained annotations, in-sufficient for depicting the detailed action dynamics required in HAU and HARn tasks. Simply combining annotations and generating captions with LLMs often lacks necessary logical and spatiotemporal consistency. In this paper, we introduce CUHK-X, a large-scale multi-modal dataset and benchmarks for HAR, HAU, and HARn. It includes 64,267 samples of 40 actions performed by 30 participants across two indoor environments, covering diverse daily scenarios. To address the challenge of spatiotemporal inconsistencies in captions, we propose a prompt-based scene creation method that leverages LLMs to generate logically connected activity sequences. CUHK-X also includes three benchmarks with six tasks to evaluate state-of-the-art models. Experimental results show average accuracies of 76.52% for HAR, 40.76% for HAU, and 70.25% for HARn. This large-scale multimodal dataset aims to empower the research community to apply, develop, and adapt data-intensive learning techniques for a wide range of human activity-related tasks.
Siyang Jiang, Mu Yuan, Bufang Yang, Lilin Xu, Yang Li 0147, Yuting He 0006, Liran Dong, Wenrui Lu, Zhenyu Yan 0002, Xiaofan Jiang 0001, Wei Gao 0006, Hongkai Chen 0001, Guoliang Xing
MobiSys14
2025 Myo-Trainer: A Vision-based Muscle-Aware Motion Feedback System for In-Home Resistance Training
abstract
In-home resistance training (RT) is a convenient and effective way to maintain health and well-being. However, incorrect exercise execution can result in unintended muscle engagement and an increased risk of injury. Without access to professional coaching, an accurate muscle-aware motion feedback system becomes essential for safe and effective training. However, existing visual language models (VLMs) struggle to provide accurate and effective muscle-aware movement guidance due to their limited understanding of RT motion and the absence of related expert knowledge. In this work, we introduce Myo-Trainer, the first vision-based muscle-aware motion feedback system that uses explicit muscle-aware motion analysis and domain-specific expert knowledge to provide corrective guidance on muscle engagement and movement execution. Also, we propose a novel DAGCN-Former network that integrates both spatial and temporal modeling capabilities to capture the complex dynamics of human RT motion. Experiments involving 26 subjects and 1000+ minutes of RT demonstrate that Myo-Trainer improves the accuracy of motion analysis by 17.22%, achieves a 2.5x reduced inference latency and a BertScore of 85.88% of generated feedback compared to those provided by experienced certified trainers, outperforming existing solutions. Additionally, Myo-Trainer received higher satisfaction ratings from participants compared to other AI trainers and video tutorials, highlighting its potential for real-world applications.
Yuting He 0006, Xinyan Wang 0003, Mu Yuan, Bufang Yang, Siyang Jiang, Yihua Huang 0002, Doris Sau-Fung Yu, Guoliang Xing, Hongkai Chen 0001
MobiCom9
2025 Poster: Mobile Menstrual Health Advising with Multimodal Feature Engineering
Liekang Zeng, Zhenyu Yan 0002, Yunqi Guo, Hongkai Chen 0001, Guoliang Xing
MobiSys5
2025 ContextAgent: Context-Aware Proactive LLM Agents with Open-world Sensory Perceptions
abstract
Recent advances in Large Language Models (LLMs) have propelled intelligent agents from reactive responses to proactive support. While promising, existing proactive agents either rely exclusively on observations from enclosed environments (e.g., desktop UIs) with direct LLM inference or employ rule-based proactive notifications, leading to suboptimal user intent understanding and limited functionality for proactive service. In this paper, we introduce ContextAgent, the first context-aware proactive agent that incorporates extensive sensory contexts surrounding humans to enhance the proactivity of LLM agents. ContextAgent first extracts multi-dimensional contexts from massive sensory perceptions on wearables (e.g., video and audio) to understand user intentions. ContextAgent then leverages the sensory contexts and personas from historical data to predict the necessity for proactive services. When proactive assistance is needed, ContextAgent further automatically calls the necessary tools to assist users unobtrusively. To evaluate this new task, we curate ContextAgentBench, the first benchmark for evaluating context-aware proactive LLM agents, covering 1,000 samples across nine daily scenarios and twenty tools. Experiments on ContextAgentBench show that ContextAgent outperforms baselines by achieving up to 8.5% and 6.0% higher accuracy in proactive predictions and tool calling, respectively. We hope our research can inspire the development of more advanced, human-centric, proactive AI assistants. The code and dataset are publicly available at https://github.com/openaiotlab/ContextAgent.
Bufang Yang, Lilin Xu, Liekang Zeng, Kaiwei Liu 0001, Siyang Jiang, Wenrui Lu, Hongkai Chen 0001, Xiaofan Jiang 0001, Guoliang Xing, Zhenyu Yan 0002
NeurIPS7
2025 SHADE-AD: An LLM-Based Framework for Synthesizing Activity Data of Alzheimer's Patients
abstract
Alzheimer's Disease (AD) has become an increasingly critical global health concern, which necessitates effective monitoring solutions in smart health applications. However, the development of such solutions is significantly hindered by the scarcity of AD-specific activity datasets. To address this challenge, we propose SHADE-AD, a Large Language Model (LLM) framework for Synthesizing Human Activity Datasets Embedded with AD features. Leveraging both public datasets and our own collected data from 99 AD patients, SHADE-AD synthesizes human activity videos that specifically represent AD-related behaviors. By employing a three-stage training mechanism, it broadens the range of activities beyond those collected from limited deployment settings. We conducted comprehensive evaluations of the generated dataset, demonstrating significant improvements in downstream tasks such as Human Activity Recognition (HAR) detection, with enhancements of up to 79.69%. Detailed motion metrics between real and synthetic data show strong alignment, validating the realism and utility of the synthesized dataset. These results underscore SHADE-AD's potential to advance smart health applications by providing a cost-effective, privacy-preserving solution for AD monitoring.
Heming Fu, Hongkai Chen 0001, Shan Lin 0001, Guoliang Xing
SenSys2
2025 Cumulative-Time Signal Temporal Logic
abstract
Signal Temporal Logic (STL) is a widely adopted specification language for Cyber-Physical Systems that can be used to express critical temporal requirements, such as system safety and response time. STL’s expressivity, however, is not sufficient to capture the cumulative duration during which a property holds within an interval of time. To overcome this limitation, we introduce Cumulative-Time Signal Temporal Logic (CT-STL) which operates over discrete-time signals and extends STL with a new cumulative-time operator. This operator compares the sum of all timesteps for which its nested formula is true with a threshold. We present both a qualitative and a quantitative (robustness) semantics for CT-STL and prove the soundness and completeness of the robustness semantics. We also provide an efficient online monitoring algorithm for both semantics. We demonstrate the utility of CT-STL via two case studies: specifying and monitoring cumulative temporal requirements for a microgrid and an artificial pancreas.
Hongkai Chen 0001, Shouvik Roy, Ezio Bartocci, Scott A. Smolka, Scott D. Stoller, Shan Lin 0001
ACM Trans. Embed. Comput. Syst.1
2024 Demo Abstract: AD-CLIP: Privacy-Preserving, Low-Cost Synthetic Human Action Dataset for Alzheimer's Patients via CLIP-based Models
abstract
With the increasing demand for smart health applications that emphasize privacy and efficiency, we introduce AD-CLIP, a synthetic data generation framework using CLIP-based models for Alzheimer’s patients. Leveraging the public dataset and data we collected from Alzheimer’s patients, AD-CLIP synthesizes human action videos featuring Alzheimer’s disease. To address privacy concerns, labeling cost, and imbalanced data distribution, AD-CLIP generates a comprehensive labeled human action skeleton dataset from depth cameras with balanced data distribution. Our preliminary experiments confirm the effectiveness of the synthesized dataset by improving the accuracy of human activity recognition up to 76.56%, which demonstrates AD-CLIP’s potential to enhance smart health applications.
Heming Fu, Hongkai Chen 0001, Guoliang Xing
IPSN2
2024 Improving On-Device LLMs' Sensory Understanding with Embedding Interpolations
abstract
Large Language Models (LLMs) have shown significant potential in performing inferences on various tasks using heterogeneous sensors with minimal human intervention. Despite their promise, challenges such as high inference overhead and limitations on resource-constrained edge devices remain. Additionally, model hallucinations, particularly those arising from cognitive biases when interpreting numerical data, hinder performance. This work introduces a novel technique, embedding interpolation, to enhance LLMs' understanding of sensor measurements and mitigate inference overhead on edge devices. By computing embeddings through pre-computed boundary embeddings instead of directly from the input, we improve efficiency and accuracy. The effective-ness of this approach is demonstrated through visualizations with image generation models.
Kaiyuan Hou, Yunqi Guo, Heming Fu, Hongkai Chen 0001, Zhenyu Yan 0002, Guoliang Xing, Xiaofan Jiang 0001
MobiCom4
2024 Demo: EmoMarker: A Privacy-Preserving, Multi-Modal Sensing System for Dyadic Digital Biomarkers of Expressed Emotions for Patients with Dementia
abstract
Alzheimer's disease and related dementia has emerged as a global health challenge due to aging population. Expressed Emotion (EE) is a widely-used medical measure of family emotional environment of patients with caregivers. We present EmoMarker, a multi-modal sensor detection system for dyadic digital biomarkers of EE in dementia patients' homes. EmoMarker consists of a privacy-preserving depth camera and a microphone to extracts interpretable dyadic (i.e., motor and acoustic) digital biomarkers of the interaction between patients and caregivers and predict the scores of Family Altitude Scale, an assessment tool for measuring the emotional climate of families. We have deployed our system in 99 elder people's homes and achieved 81.13% prediction accuracy in preliminary results.
Yang Li 0147, Doris Sau-Fung Yu, Shuangzhou Chen, Guoliang Xing, Hongkai Chen 0001
MobiSys5
2024 Demo: MuRa: A Scalable Mobile Ultra-wideband Testbed for Multi-node Ranging
abstract
Ultra-wideband (UWB) technology, known for its precise distance measurement capabilities, is widely utilized in indoor localization and ranging applications. However, recent UWB systems that operate within dense and dynamic networks face challenges in real-time performance evaluation and experimental data collection, especially in infrastructure-less scenarios. To address this, we present a mobile UWB testbed, which consists of a scalable ranging system, a data collection pipeline, and a mesh-based control network. This testbed facilitates UWB-related research such as ranging protocol design, as well as implementing and evaluating applications such as social interaction analysis and contact tracing.
Shaoyang Yang, Fang Liu 0022, Guoliang Xing, Hongkai Chen 0001
MobiSys5
2024 Demo: Myotrainer: Muscle-Aware Motion Analysis and Feedback System for In-Home Resistance Training
abstract
Resistance training is widely incorporated in exercise programs, including in-home fitness and rehabilitation. However, improper motion patterns and muscle stimulation can undermine the safety of the subjects, making precise monitoring essential. Existing solutions primarily focus on correcting motion patterns with difficulties assessing muscle contraction levels. In this work, we introduce MyoTrainer, which provides muscle-aware motion descriptions and personalized feedback in natural language. Taking a person's exercise video as input, MyoTrainer first utilizes pose estimation models to capture motion sequences in real-time. A GCN-Former model has been developed for fine-grained motion analysis, which includes action recognition, incorrect movement pattern detection, and muscle contraction intensity estimation. Additionally, MyoTrainer integrates fitness and physiotherapeutic domain knowledge to deliver personalized, professional feedback. Extensive evaluations show that our system outperforms existing solutions in all recognition tasks and a survey indicates 88.9% of users find the generated feedback to be beneficial.
Yuting He 0006, Xinyan Wang 0003, Mu Yuan, Di Duan, Doris Sau-Fung Yu, Guoliang Xing, Hongkai Chen 0001
SenSys7
2023 An STL-based Approach to Resilient Control for Cyber-Physical Systems
abstract
We present ResilienC, a framework for resilient control of Cyber-Physical Systems subject to STL-based requirements. ResilienC utilizes a recently developed formalism for specifying CPS resiliency in terms of sets of (rec, dur) real-valued pairs, where rec represents the system’s capability to rapidly recover from a property violation (recoverability), and dur is reflective of its ability to avoid violations post-recovery (durability). We define the resilient STL control problem as one of multi-objective optimization, where the recoverability and durability of the desired STL specification are maximized. When neither objective is prioritized over the other, the solution to the problem is a set of Pareto-optimal system trajectories. We present a precise solution method to the resilient STL control problem using a mixed-integer linear programming encoding and an a posteriori ϵ -constraint approach for efficiently retrieving the complete set of optimally resilient solutions. In ResilienC, at each time-step, the optimal control action selected from the set of Pareto-optimal solutions by a Decision Maker strategy realizes a form of Model Predictive Control. We demonstrate the practical utility of the ResilienC framework on two significant case studies: autonomous vehicle lane keeping and deadline-driven, multi-region package delivery.
Hongkai Chen 0001, Scott A. Smolka, Nicola Paoletti, Shan Lin 0001
HSCC1
2020 Data-Driven Robust Control for a Closed-Loop Artificial Pancreas
abstract
We present a fully closed-loop design for an artificial pancreas (AP) that regulates the delivery of insulin for the control of Type I diabetes. Our AP controller operates in a fully automated fashion, without requiring any manual interaction with the patient (e.g., in the form of meal announcements). A major obstacle to achieving closed-loop insulin control are the "unknown disturbances" related to various aspects of a patient's daily behavior, especially meals and physical activity. Such disturbances can significantly affect the patient's blood glucose levels. To handle such uncertainties, we present a data-driven, robust, model-predictive control framework in which we capture a wide range of individual meal and exercise patterns using uncertainty sets learned from historical data. These uncertainty sets are then used in the insulin controller to achieve automated, precise, and personalized insulin therapy. We provide an extensive in silico evaluation of our robust AP design, demonstrating the potential of the approach. In particular, without the benefit of explicit meal announcements, our approach can regulate glucose levels for large clusters of meal profiles learned from population-wide survey data and cohorts of virtual patients, even in the presence of high carbohydrate disturbances.
Nicola Paoletti, Kin Sum Liu, Hongkai Chen 0001, Scott A. Smolka, Shan Lin 0001
IEEE ACM Trans. Comput. Biol. Bioinform.3
2019 MagTrack: Enabling Safe Driving Monitoring with Wearable Magnetics
abstract
"Hands on the wheel, eyes on the road" is the central guideline of safe vehicle driving practices. Many advanced driver assistance systems can effectively detect abnormal vehicle motions. However, these systems often leave insufficient time for drivers to respond to complex road situations, especially when the drivers are distracted. To reduce accidents, it is essential to detect whether a driver complies with safe driving guidelines in real time and provide warnings early before any dangerous maneuvers occur. There are vision-based driver distraction monitoring systems which rely on cameras in high-end vehicles, but their performances are heavily constrained by visibility requirements. In this paper, we present MagTrack, a driver monitoring system that is based on tracking magnetic tags worn by the user. With a single smartwatch and two low-cost magnetic accessories: a hand magnetic ring and a head magnetic eyeglasses clip, our system tracks and classifies a driver's bimanual and head movements simultaneously using both analytical and approximation sensing models. Our approach is robust to driver's postures, vehicles, and environmental changes. We demonstrate that a wide range of activities can be detected by our system, including bimanual steering, visual and manual distractions, and lane changes and turns. In extensive road tests with 500+ instances of driving activities and 500+ minutes of road driving with 10 subjects, MagTrack achieves 87% of precision and 90% of recall rate on the detection of unsafe driving activities.
Hua Huang 0003, Hongkai Chen 0001, Shan Lin 0001
MobiSys2