Heming Fu

dblp:174/0541 · DBLP profile ↗
← Back
8ranked-venue papers
3as first author
7since 2021 · last 2026
0000-0002-2221-0143ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 6 · 2 first-author · 6 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 MultiFinBen: Benchmarking Large Language Models for Multilingual and Multimodal Financial Application
abstract
Xueqing Peng, Lingfei Qian, Yan Wang, Ruoyu Xiang, Yueru He, Yang Ren, Mingyang Jiang, Vincent Jim Zhang, Yuqing Guo, Jeff Zhao, Huan He, Yi Han, Yun Feng, Yuechen Jiang, Yupeng Cao, Haohang Li, Yangyang Yu, Xiaoyu Wang, Penglei Gao, Shengyuan Lin, Keyi Wang, Shanshan Yang, Yilun Zhao, Zhiwei Liu, Peng Lu, Jerry Huang, Suyuchen Wang, Triantafillos Papadopoulos, Polydoros Giannouris, Efstathia Soufleri, Nuo Chen, Zhiyang Deng, Heming Fu, Yijia Zhao, Mingquan Lin, Meikang Qiu, Kaleb E Smith, Arman Cohan, Xiao-Yang Liu, Jimin Huang, Guojun Xiong, Alejandro Lopez-Lira, Xi Chen, Junichi Tsujii, Jian-Yun Nie, Sophia Ananiadou, Qianqian Xie. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Xueqing Peng, Lingfei Qian, Yan Wang 0015, Ruoyu Xiang, Yueru He, Mingyang Jiang, Vincent Jim Zhang, Jeff Zhao, Yuechen Jiang, Yupeng Cao, Haohang Li, Yangyang Yu, Penglei Gao, Shengyuan Lin, Yilun Zhao 0001, Zhiwei Liu 0003, Peng Lu 0006, Jerry Huang, Suyuchen Wang, Triantafillos Papadopoulos, Polydoros Giannouris, Efstathia Soufleri, Nuo Chen 0002, Zhiyang Deng, Heming Fu, Yijia Zhao, Mingquan Lin, Meikang Qiu, Kaleb E. Smith, Arman Cohan, Xiao-Yang Liu, Jimin Huang, Guojun Xiong, Alejandro Lopez-Lira, Xi Chen 0003, Jun'ichi Tsujii, Jian-Yun Nie, Sophia Ananiadou, Qianqian Xie
ACL (1)33
2025 SHADE-AD: An LLM-Based Framework for Synthesizing Activity Data of Alzheimer's Patients
abstract
Alzheimer's Disease (AD) has become an increasingly critical global health concern, which necessitates effective monitoring solutions in smart health applications. However, the development of such solutions is significantly hindered by the scarcity of AD-specific activity datasets. To address this challenge, we propose SHADE-AD, a Large Language Model (LLM) framework for Synthesizing Human Activity Datasets Embedded with AD features. Leveraging both public datasets and our own collected data from 99 AD patients, SHADE-AD synthesizes human activity videos that specifically represent AD-related behaviors. By employing a three-stage training mechanism, it broadens the range of activities beyond those collected from limited deployment settings. We conducted comprehensive evaluations of the generated dataset, demonstrating significant improvements in downstream tasks such as Human Activity Recognition (HAR) detection, with enhancements of up to 79.69%. Detailed motion metrics between real and synthetic data show strong alignment, validating the realism and utility of the synthesized dataset. These results underscore SHADE-AD's potential to advance smart health applications by providing a cost-effective, privacy-preserving solution for AD monitoring.
Heming Fu, Hongkai Chen 0001, Shan Lin 0001, Guoliang Xing
SenSys1
2024 Demo Abstract: AD-CLIP: Privacy-Preserving, Low-Cost Synthetic Human Action Dataset for Alzheimer's Patients via CLIP-based Models
abstract
With the increasing demand for smart health applications that emphasize privacy and efficiency, we introduce AD-CLIP, a synthetic data generation framework using CLIP-based models for Alzheimer’s patients. Leveraging the public dataset and data we collected from Alzheimer’s patients, AD-CLIP synthesizes human action videos featuring Alzheimer’s disease. To address privacy concerns, labeling cost, and imbalanced data distribution, AD-CLIP generates a comprehensive labeled human action skeleton dataset from depth cameras with balanced data distribution. Our preliminary experiments confirm the effectiveness of the synthesized dataset by improving the accuracy of human activity recognition up to 76.56%, which demonstrates AD-CLIP’s potential to enhance smart health applications.
Heming Fu, Hongkai Chen 0001, Guoliang Xing
IPSN1
2024 Improving On-Device LLMs' Sensory Understanding with Embedding Interpolations
abstract
Large Language Models (LLMs) have shown significant potential in performing inferences on various tasks using heterogeneous sensors with minimal human intervention. Despite their promise, challenges such as high inference overhead and limitations on resource-constrained edge devices remain. Additionally, model hallucinations, particularly those arising from cognitive biases when interpreting numerical data, hinder performance. This work introduces a novel technique, embedding interpolation, to enhance LLMs' understanding of sensor measurements and mitigate inference overhead on edge devices. By computing embeddings through pre-computed boundary embeddings instead of directly from the input, we improve efficiency and accuracy. The effective-ness of this approach is demonstrated through visualizations with image generation models.
Kaiyuan Hou, Yunqi Guo, Heming Fu, Hongkai Chen 0001, Zhenyu Yan 0002, Guoliang Xing, Xiaofan Jiang 0001
MobiCom3
2024 ADMarker: A Multi-Modal Federated Learning System for Monitoring Digital Biomarkers of Alzheimer's Disease
abstract
Alzheimer's Disease (AD) and related dementia are a growing global health challenge due to the aging population. In this paper, we present ADMarker, the first end-to-end system that integrates multi-modal sensors and new federated learning algorithms for detecting multidimensional AD digital biomarkers in natural living environments. ADMarker features a novel three-stage multi-modal federated learning architecture that can accurately detect digital biomarkers in a privacy-preserving manner. Our approach collectively addresses several major real-world challenges, such as limited data labels, data heterogeneity, and limited computing resources. We built a compact multi-modality hardware system and deployed it in a four-week clinical trial involving 91 elderly participants. The results indicate that ADMarker can accurately detect a comprehensive set of digital biomarkers with up to 93.8% accuracy and identify early AD with an average of 88.9% accuracy. ADMarker offers a new platform that can allow AD clinicians to characterize and track the complex correlation between multidimensional interpretable digital biomarkers, demographic factors of patients, and AD diagnosis in a longitudinal manner.
Xiaomin Ouyang, Xian Shuai, Yang Li 0147, Li Pan 0004, Xifan Zhang, Heming Fu, Sitong Cheng, Xinyan Wang 0003, Shihua Cao, Jiang Xin, Hazel Mok, Zhenyu Yan 0002, Doris Sau-Fung Yu, Timothy Kwok, Guoliang Xing
MobiCom6
2023 Harmony: Heterogeneous Multi-Modal Federated Learning through Disentangled Model Training
abstract
Multi-modal sensing systems are increasingly prevalent in real-world applications such as health monitoring and autonomous driving. Most multi-modal learning approaches need to access users' raw data, which poses significant concerns to users' privacy. Federated learning (FL) provides a privacy-aware distributed learning framework. However, current FL approaches have not addressed the unique challenges of heterogeneous multi-modal FL systems, such as modality heterogeneity and significantly longer training delay. In this paper, we propose Harmony, a new system for heterogeneous multi-modal federated learning. Harmony disentangles the multi-modal network training in a novel two-stage framework, namely modality-wise federated learning and federated fusion learning. By integrating a novel balance-aware resource allocation mechanism in modality-wise FL and exploiting modality biases in federated fusion learning, Harmony improves the model accuracy under non-i.i.d. data distributions and speeds up system convergence. We implemented Harmony on a real-world multi-modal sensor testbed deployed in the homes of 16 elderly subjects for Alzheimer's Disease monitoring. Our evaluation on the testbed and three large-scale public datasets of different applications show that, Harmony outperforms by up to 46.35% accuracy over state-of-the-art baselines and saves up to 30% training delay.
Xiaomin Ouyang, Heming Fu, Sitong Cheng, Li Pan 0004, Neiwen Ling, Guoliang Xing, Jianwei Huang 0001
MobiSys3
2022 An Indoor Smart Traffic Dataset and Data Collection System: Dataset
abstract
Smart traffic is an emerging research area gaining more attention due to a class of emerging applications such as autonomous driving. Most smart traffic scenarios are outdoors, which are hard to collect traffic data and build demanding sensing systems. In this work, an indoor smart traffic testbed with an F1TENTH autonomous driving vehicle is built, allowing the collection of traffic datasets under different scenarios and performing various smart traffic tasks. This novel data collection system and collected dataset can help research teams build various smart traffic systems and evaluate indoor smart traffic datasets. The collected traffic light dataset is publicly available at the link1.
Neiwen Ling, Nan Guan, Heming Fu, Guoliang Xing
SenSys4
2016 Fast detection of impact location using kernel extreme learning machine
Heming Fu, Chi-Man Vong, Pak-Kin Wong 0001, Zhi-Xin Yang 0001
Neural Comput. Appl.1