VLDB 2026 Research / reviewers in the wild / expert
Xiaomin Ouyang
dblp:295/6332
· DBLP profile ↗
22ranked-venue papers
7as first author
22since 2021 · last 2026
0000-0003-0710-0963ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 19 · 7 first-author · 19 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Vistar: Enhancing the Perception Capability of LLMs under Imprecise IMU-Text AlignmentabstractThis paper introduces Vistar, a novel self-supervised framework for inertial measurement unit (IMU) signal perception designed for large language models (LLMs). Unlike visual data, IMU signals are high-frequency time series with low interpretability, making manual annotation with natural language particularly challenging. Even when using vision-language models (VLMs) to describe events in videos synchronized with IMU signals, a semantic gap remains between high-level visual semantics and low-level IMU vibrations. The core idea of Vistar is to achieve accurate IMU signal perception through collaborations between offline cross-modal alignment and online retrieval-augmented generation. During offline training, Vistar uses pretrained vision and language encoders as anchors to learn IMU encoders via hierarchical cross-modal contrastive learning, establishing both inter- and intra-sample alignment. Given that the enhanced training strategy still fails to achieve precise alignment between IMU and text, during online inference, Vistar further employs a retrieval-augmented generation mechanism to generate distilled textual descriptions from similar text filtered based on structural relations of their paired IMU samples. Extensive evaluations on three multimodal datasets demonstrate that Vistar consistently outperforms state-of-the-art (SOTA) baselines by up to 57.45% in IMU-to-text retrieval and improves the generated text similarity with ground truths in IMU perception by up to 31.90%. Yatong Chen 0001, Chenzhi Hu, Ruijie Wang 0004, Xiaomin Ouyang, Shengzhong Liu, Jianxin Li 0002, Fan Wu 0006, Guihai Chen |
KDD (1) | 5 |
| 2026 | FBLayout: Optimizing Memory Layout for Efficient LLM Finetuning on Mobile GPUsabstractTransformer-based models have enabled unprecedented capabilities across language, vision, and multimodal tasks. On-device fine-tuning of transformer models offers a privacy-preserving path to personalized AI, yet remains inefficient on mobile GPUs due to severe memory constraints and frequent layout transformations in attention mechanism during training. Existing mobile training frameworks either use unified layouts for forward and backward passes - leading to fragmented memory access and poor GPU utilization during backpropagation - or rely on explicit layout conversions, which introduce significant transformation overhead.To overcome this, we propose FBLayout, a layout-aware framework that co-designs tensor organization with mobile GPU platforms. FBLayout introduces: (1) a unified R-Tile layout for multidimensional reductions across forward/backward passes; (2) tile-based index transformation to eliminate physical data movement; and (3) activation-guided layout selection to propagate efficient layouts globally. Evaluations on seven transformer models across different mobile phones (including ARM Mali and Qualcomm Adreno GPUs) show that FBLayout achieves 2.2-5.7× speedup over MNN, TFLite, and TVM, while significantly improving cache efficiency and reducing memory footprint, enabling practical on-device large model fine-tuning. Kahou Tam, Wei Niu 0002, Xiaomin Ouyang, Cheng-Zhong Xu 0001, Li Li 0064 |
MobiSys | 4 |
| 2026 | MMEdge: Accelerating On-device Multimodal Inference via Pipelined Sensing and EncodingabstractReal-time multimodal inference on resource-constrained edge devices is essential for applications such as autonomous driving, human-computer interaction, and mobile health. However, prior work often overlooks the tight coupling between sensing dynamics and model execution, as well as the complex inter-modality dependencies. In this paper, we propose MMEdge, a new on-device multimodal inference framework based on pipelined sensing and encoding. Instead of waiting for complete sensor inputs, MMEdge decomposes the entire inference process into a sequence of fine-grained sensing and encoding units, allowing computation to proceed incrementally as data arrive. MMEdge also introduces a lightweight but effective temporal aggregation module that captures rich temporal dynamics across different pipelined units to maintain accuracy performance. Such pipelined design also opens up opportunities for fine-grained cross-modal optimization and early decision-making during inference. To further enhance system performance under resource variability and input data complexity, MMEdge incorporates an adaptive multimodal configuration optimizer that dynamically selects optimal sensing and model configurations for each modality under latency constraints, and a cross-modal speculative skipping mechanism that bypasses future units of slower modalities when early predictions reach sufficient confidence. We evaluate MMEdge using two public multimodal datasets and deploy it on a real-world unmanned aerial vehicle (UAV)-based multimodal testbed. The results show that MMEdge significantly reduces end-to-end latency while maintaining high task accuracy across various system and data dynamics. A video demonstration of MMEdge’s performance in real world is available at https://youtu.be/qRew7sT-iWw. Runxi Huang, Mingxuan Yu, Mingyu Tsoi, Xiaomin Ouyang |
SenSys | 4 |
| 2026 | Short Paper: WearBCI Dataset: Understanding and Benchmarking Real-World Wearable Brain-Computer Interfaces SignalsabstractBrain-computer interfaces (BCIs) have opened new platforms for human-computer interaction, medical diagnostics, and neurorehabilitation. Wearable BCI systems, which typically employ non-invasive electrodes for portable monitoring, hold great promise for real-world applications, but also face significant challenges of signal quality degradation caused by motion artifacts and environmental interferences. Most existing wearable BCI datasets are collected under stationary or controlled lab settings, limiting their utility for evaluating performance under body movement. To bridge this gap, we introduce WearBCI, the first dataset that comprehensively evaluates wearable BCI signals under different motion dynamics with synchronized multimodal recordings (EEG, IMU, and egocentric video), and systematic benchmark evaluations for studying impacts of motion artifact. Specifically, we collect data from 36 participants across different motion dynamics, including body movements, walking, and navigation. This dataset includes synchronized electroencephalography (EEG), inertial measurement unit (IMU) data, and egocentric video recordings. We analyze the collected wearable EEG signals to understand the impact of motion artifacts across different conditions, and benchmark representative EEG signal enhancement techniques on our dataset. Furthermore, we explore two new case studies: cross-modal EEG signal enhancement and multi-dimension human behavior understanding. These findings offer valuable insights into real-world wearable BCI deployment and new applications. Haoxian Liu, Hengle Jiang, Lanxuan Hong, Xiaomin Ouyang |
SenSys | 4 |
| 2026 | MoViD: View-Invariant 3D Human Pose Estimation via Motion-View Disentanglementabstract3D human pose estimation is a key enabling technology for applications such as healthcare monitoring, human-robot collaboration, and immersive gaming, but real-world deployment remains challenged by viewpoint variations. Existing methods struggle to generalize to unseen camera viewpoints, require large amounts of training data, and suffer from high inference latency. We propose MoViD, a viewpoint-invariant 3D human pose estimation framework that disentangles viewpoint information from motion features. The key idea is to extract viewpoint information from intermediate pose features and leverage it to enhance both the robustness and efficiency of pose estimation. MoViD introduces a view estimator that models key joint relationships to predict viewpoint information, and an orthogonal projection module to disentangle motion and view features, further enhanced through physics-grounded contrastive alignment across views. For real-time edge deployment, MoViD employs a frame-by-frame inference pipeline with a view-aware strategy that adaptively activates flip refinement based on the estimated viewpoint. Evaluations on nine public datasets and newly collected multiview UAV and gait analysis datasets show that MoViD reduces pose estimation error by over 24.2% compared to state-of-the-art methods, maintains robust performance under severe occlusions with 60% less training data, and achieves real-time inference at 15 FPS on NVIDIA edge devices. Yejia Liu, Hengle Jiang, Haoxian Liu, Runxi Huang, Xiaomin Ouyang |
SenSys | 5 |
| 2026 | Can LLMs Be Effective Sensor Processing Copilots?abstractEffective sensor data processing is critical for cyber-physical and IoT systems but often requires specialized expertise. While Large Language Models (LLMs) show promise as autonomouscopilotsfor sensor processing, their capabilities remain underexplored. We introduce SensorBench, the first comprehensive benchmark for evaluating LLMs across diverse real-world sensor datasets and tasks. SensorBench evaluates three paradigms for leveraging LLMs in sensing tasks: Tool-Augmented Coding (TAC), Standalone Coding (SAC), and Direct Answer (DA). We evaluate 8 leading LLM variants, including 2 Large Reasoning Models (LRMs) and 2 domain-specific LLMs, providing a structured reference for absolute performance, latency, and resource requirements. Our analysis reveals that: (1) TAC significantly outperforms SAC and DA; (2) LLMs excel at simple tasks but consistently underperform domain experts on compositional tasks requiring parameter tuning and multi-step reasoning. (3) The reasoning mechanism introduced in LRMs does not yield substantial performance gains. To improve the performance, we explore four prompting strategies and fine-tuning approaches (using our newly released sensor-processing corpus). The results show that self-verification prompting proves most effective, outperforming other methods simultaneously in 48% of tasks, while fine-tuning yields marginal gains. Our analysis suggests that more sophisticated interaction frameworks, such as signal-level self-verification, may bridge the gap to human expert-level performance. This benchmark provides a foundation for evaluating and improving LLMs in sensing applications1. Pengrui Quan, Xiaomin Ouyang, Jeya Vikranth Jeyakumar, Ziqi Wang 0001, Yang Xing 0003, Mani Srivastava 0001 |
IEEE Internet Things J. | 2 |
| 2026 | COOL: A Cloud-Fog Federated Learning System With Multimodal Isolated Client Data in Edge NetworksabstractCloud-fog Federated Learning (FL) is promising for collaborative model training in large-scale edge networks. In cloud-fog FL with constrained communication resources, selecting high-quality client models for aggregation is critical to boost the global model. However, the clients in real world hold multimodal and heterogeneous data, while existing selection strategies rarely consider the imbalance of communication cost and model convergence rates across modalities, thus seriously degrading the efficiency of model aggregation. Moreover, most previous multimodal fusion methods require aligned multimodal samples. However, the data of modality-heterogeneous clients may be isolated and unaligned in FL, so these previous methods cannot be applied to such scenarios, thereby hindering the knowledge fusion across various modalities. To address the above issues, we propose a cloud-fog FL system named COOL, which achievesunimodal aggregationat the fog layer andmultimodal fusionat the cloud layer. First, we propose a Modality-aware Online Client Selection (MOCS) strategy to assist unimodal aggregation. Unlike previous selection strategies, MOCS realizes the dynamic selection budget allocation for various modalities by monitoring the convergence gap of modalities, thus striking the performance balance among modality-heterogeneous clients. Second, to overcome the limitation of previous methods that require aligned multimodal data, we propose a Multimodal Fusion strategy with Feature Synthesis (MFFS). MFFS realizes multimodal fusion with isolated samples via adaptive feature synthesis and cross-modal attention training, thus building a more powerful multimodal predictor while preserving the client data privacy. Finally, experimental results demonstrate that COOL has superior performance compared to the existing algorithms. Jialin Guo, Jianheng Tang 0001, Anfeng Liu, Naixue Xiong, Jie Wu 0001, Xiaomin Ouyang, Jun Huang 0001 |
IEEE Trans. Mob. Comput. | 6 |
| 2025 | Demo: FreePose: Real-Time View-Invariant 3D Human Pose Estimation via Motion-View Disentanglementabstract3D human pose estimation is a key technology for applications like healthcare and robotics, but its performance in real-world deployments is often compromised by viewpoint variations. We propose FreePose, a novel framework that achieves viewpoint invariance by explicitly disentangling motion and view features. The core of FreePose is a lightweight view estimator that predicts camera viewpoint from intermediate pose features. This information is then used to guide robust feature alignment and enable a view-aware inference pipeline that adaptively optimizes for latency on edge devices. We will demonstrate our system using a single Intel RealSense D435 camera, capturing from varying viewpoints throughout the demo, with real-time pose inference performed on a PC and an NVIDIA Jetson Orin NX. By leveraging FreePose, our framework achieves consistent accuracy and high frame rates across varying camera angles and positions. A video demonstration of FreePose's performance is available at https://youtu.be/tDIpCcbaRXc. Yejia Liu, Hengle Jiang, Xiaomin Ouyang |
MobiCom | 3 |
| 2025 | MMBind: Unleashing the Potential of Distributed and Heterogeneous Data for Multimodal Learning in IoTabstractMultimodal sensing systems are increasingly prevalent in various real-world applications. Most existing multimodal learning approaches heavily rely on training with a large amount of synchronized, complete multimodal data. However, such a setting is impractical in real-world IoT sensing applications where data is typically collected by distributed nodes with heterogeneous data modalities, and is also rarely labeled. In this paper, we propose MMBind, a new data binding approach for multimodal learning on distributed and heterogeneous IoT data. The key idea of MMBind is to construct a pseudo-paired multimodal dataset for model training by binding data from disparate sources and incomplete modalities through a sufficiently descriptive shared modality. We also propose a weighted contrastive learning approach to handle domain shifts among disparate data, coupled with an adaptive multimodal learning architecture capable of training models with heterogeneous modality combinations. Evaluations on ten real-world multi-modal datasets highlight that MMBind outperforms state-of-the-art baselines under varying degrees of data incompleteness and domain shift, and holds promise for advancing multimodal foundation model training in IoT applications1. Xiaomin Ouyang, Tomoyoshi Kimura, Gunjan Verma, Tarek F. Abdelzaher, Mani Srivastava 0001 |
SenSys | 1 |
| 2025 | Demo Abstract: An LLM-Powered Multimodal Mobile Sensing System for Personalized and Interactive Health Behavior AnalysisabstractCharacterizing human behaviors using mobile devices is crucial for the longitudinal monitoring of chronic diseases, such as mental health conditions and Alzheimer's Disease. Current solutions often either focus on detecting basic human activities (e.g., sitting, walking) or lack comprehensive analysis and interactive guidance based on sensor data. We present MobiBox, a lightweight mobile app for long-term behavior data collection and interactive health analysis. MobiBox captures multimodal data including high-resolution 9-axis IMU data and contextual information such as APP usage and network activities, and integrates Large Language Models (LLMs) to generate personalized guidances like interventions and daily summaries. Moreover, MobiBox features a closed-loop design that allows users rate these guidances, building a high-quality dataset to enhance performance of LLMs on mobile health applications. The demo video is available at https://youtube.com/shorts/XgOXFoRaFIw?feature=share. Liyu Zhang 0005, Wenjie Du 0004, Kwun Ho Liu, Xiaomin Ouyang |
SenSys | 5 |
| 2025 | InfoMAE: Pair-Efficient Cross-Modal Alignment for Multimodal Time-Series Sensing SignalsabstractStandard multimodal self-supervised learning (SSL) algorithms regard cross-modal synchronization as implicit supervisory labels during pretraining, thus posing high requirements on the scale and quality of multimodal samples. These constraints significantly limit the performance of sensing intelligence in IoT applications, as the heterogeneity and the non-interpretability of time-series signals result in abundant unimodal data but scarce high-quality multimodal pairs. This paper proposes InfoMAE, a cross-modal alignment framework that tackles the challenge of multimodal pair efficiency under the SSL setting by facilitating efficient cross-modal alignment of pretrained unimodal representations. InfoMAE achieves efficient cross-modal alignment with limited data pairs through a novel information theory-inspired formulation that simultaneously addresses distribution-level and instance-level alignment. Extensive experiments on two real-world IoT applications are performed to evaluate InfoMAE's pairing efficiency to bridge pretrained unimodal models into a cohesive joint multimodal model. InfoMAE enhances downstream multimodal tasks by over 60% with significantly improved multimodal pairing efficiency. It also improves unimodal task accuracy by an average of 22%. Tomoyoshi Kimura, Osama A. Hanna, Yatong Chen 0001, Yizhuo Chen, Denizhan Kara, Tianshi Wang 0002, Jinyang Li 0004, Xiaomin Ouyang, Shengzhong Liu, Mani Srivastava 0001, Suhas N. Diggavi, Tarek F. Abdelzaher |
WWW | 9 |
| 2024 | FlexLoc: Conditional Neural Networks for Zero-Shot Sensor Perspective Invariance in Object Localization with Distributed Multimodal SensorsabstractLocalization is a critical technology for various applications ranging from navigation and surveillance to assisted living. Localization systems typically fuse information from sensors viewing the scene from different perspectives to estimate the target location while also employing multiple modalities for enhanced robustness and accuracy. Recently, such systems have employed end-to-end deep neural models trained on large datasets due to their superior performance and ability to handle data from diverse sensor modalities. However, such neural models are often trained on data collected from a particular set of sensor poses (i.e., locations and orientations). During real-world deployments, slight deviations from these sensor poses can result in extreme inaccuracies. To address this challenge, we introduce FlexLoc, which employs conditional neural networks to inject node perspective information to adapt the localization pipeline. Specifically, a small subset of model weights are derived from node poses at run time, enabling accurate generalization to unseen perspectives with minimal additional overhead. Our evaluations on a multimodal, multi-view indoor tracking dataset showcase that FlexLoc improves the localization accuracy by almost 50% in the zero-shot case (no calibration data available) compared to the baselines. The source code of FlexLoc is available in https://github.com/nesl/FlexLoc. Ziqi Wang 0001, Xiaomin Ouyang, Ho Lyun Jeong, Colin Samplawski, Lance M. Kaplan, Benjamin M. Marlin, Mani Srivastava 0001 |
IROS | 3 |
| 2024 | ADMarker: A Multi-Modal Federated Learning System for Monitoring Digital Biomarkers of Alzheimer's DiseaseabstractAlzheimer's Disease (AD) and related dementia are a growing global health challenge due to the aging population. In this paper, we present ADMarker, the first end-to-end system that integrates multi-modal sensors and new federated learning algorithms for detecting multidimensional AD digital biomarkers in natural living environments. ADMarker features a novel three-stage multi-modal federated learning architecture that can accurately detect digital biomarkers in a privacy-preserving manner. Our approach collectively addresses several major real-world challenges, such as limited data labels, data heterogeneity, and limited computing resources. We built a compact multi-modality hardware system and deployed it in a four-week clinical trial involving 91 elderly participants. The results indicate that ADMarker can accurately detect a comprehensive set of digital biomarkers with up to 93.8% accuracy and identify early AD with an average of 88.9% accuracy. ADMarker offers a new platform that can allow AD clinicians to characterize and track the complex correlation between multidimensional interpretable digital biomarkers, demographic factors of patients, and AD diagnosis in a longitudinal manner. Xiaomin Ouyang, Xian Shuai, Yang Li 0147, Li Pan 0004, Xifan Zhang, Heming Fu, Sitong Cheng, Xinyan Wang 0003, Shihua Cao, Jiang Xin, Hazel Mok, Zhenyu Yan 0002, Doris Sau-Fung Yu, Timothy Kwok, Guoliang Xing |
MobiCom | 1 |
| 2023 | Design and Deployment of Multi-Modal Federated Learning Systems for Alzheimer's Disease MonitoringabstractAlzheimer's Disease (AD) and related dementia are a growing global health challenge due to the aging population. The prominence of mobile devices and recent breakthroughs in machine learning have enabled an emerging class of new AI-powered health systems for applications like Alzheimer's Disease monitoring. In this paper, we present the first end-to-end system that integrates multi-modal sensors and federated learning algorithms for detecting multidimensional AD digital biomarkers in natural living environments. We recognize several major challenges in designing such a real-world federated learning system, including limited data labels, data heterogeneity, and limited computing resources. We built a compact multi-modality hardware system and deployed it in a four-week clinical trial involving 61 elderly participants. The results indicate that our system can accurately detect a comprehensive set of digital biomarkers with up to 95% accuracy and identify AD with an average of 87.5% accuracy. Xiaomin Ouyang |
MobiSys | 1 |
| 2023 | Harmony: Heterogeneous Multi-Modal Federated Learning through Disentangled Model TrainingabstractMulti-modal sensing systems are increasingly prevalent in real-world applications such as health monitoring and autonomous driving. Most multi-modal learning approaches need to access users' raw data, which poses significant concerns to users' privacy. Federated learning (FL) provides a privacy-aware distributed learning framework. However, current FL approaches have not addressed the unique challenges of heterogeneous multi-modal FL systems, such as modality heterogeneity and significantly longer training delay. In this paper, we propose Harmony, a new system for heterogeneous multi-modal federated learning. Harmony disentangles the multi-modal network training in a novel two-stage framework, namely modality-wise federated learning and federated fusion learning. By integrating a novel balance-aware resource allocation mechanism in modality-wise FL and exploiting modality biases in federated fusion learning, Harmony improves the model accuracy under non-i.i.d. data distributions and speeds up system convergence. We implemented Harmony on a real-world multi-modal sensor testbed deployed in the homes of 16 elderly subjects for Alzheimer's Disease monitoring. Our evaluation on the testbed and three large-scale public datasets of different applications show that, Harmony outperforms by up to 46.35% accuracy over state-of-the-art baselines and saves up to 30% training delay. Xiaomin Ouyang, Heming Fu, Sitong Cheng, Li Pan 0004, Neiwen Ling, Guoliang Xing, Jianwei Huang 0001 |
MobiSys | 1 |
| 2023 | Mozart: A Mobile ToF System for Sensing in the Dark through Phase ManipulationabstractSensing in low-light and dark environments has a wide range of applications. However, existing sensing technologies suffer several major challenges, such as excessive noise and low resolution. This paper proposes Mozart - a new mobile sensing system that leverages off-the-shelf Time-of-Flight (ToF) depth cameras to generate high-resolution and rich-in-texture maps for applications in dark scenarios. The design of Mozart is based on our key observation that the phase components of ToF measurements can be manipulated to expose texture information. Through in-depth analysis of the physical reflection model, we show that the textures can be exposed and enhanced using highly compute-efficient phase manipulation functions. By exploiting the physics texture models, we propose an autoencoder-based unsupervised learning approach that can automatically learn efficient representations from phase components to generate high-resolution maps. We implemented Mozart on several Android smartphone models1, and an edge testbed with standalone ToF camera platforms for various applications in the dark. The results show that Mozart can work in real time and delivers significant improvement over existing sensing technologies. Therefore, Mozart offers a low-cost, high-performance sensing technology for next-generation applications in the dark. Xiaomin Ouyang, Li Pan 0004, Wenrui Lu, Guoliang Xing, Xiaoming Liu 0002 |
MobiSys | 2 |
| 2023 | ClusterFL: A Clustering-based Federated Learning System for Human Activity RecognitionabstractFederated Learning (FL) has recently received significant interest, thanks to its capability of protecting data privacy. However, existing FL paradigms yield unsatisfactory performance for a wide class of human activity recognition (HAR) applications, since they are oblivious to the intrinsic relationship between data of different users. We propose ClusterFL, a clustering-based federated learning system that can provide high model accuracy and low communication overhead for HAR applications. ClusterFL features a novel clustered multi-task federated learning framework that minimizes the empirical training loss of multiple learned models while automatically capturing the intrinsic clustering relationship among the nodes. We theoretically prove the convergence of proposed FL framework for non-convex and strongly convex models and provide the guidance on selection of hyper-parameters for achieving such convergence. Based on the learned cluster relationship, ClusterFL can efficiently drop the nodes that converge slower or have little correlations with others in each cluster, significantly speeding up the convergence while maintaining the accuracy performance. We evaluate the performance of ClusterFL on an NVIDIA edge testbed using four new HAR datasets collected from 145 users. The results show that ClusterFL outperforms several state-of-the-art FL paradigms in terms of overall accuracy and can save more than 50% communication overhead. Xiaomin Ouyang, Guoliang Xing, Jianwei Huang 0001 |
ACM Trans. Sens. Networks | 1 |
| 2022 | Cosmo: contrastive fusion learning with small data for multimodal human activity recognitionabstractHuman activity recognition (HAR) is a key enabling technology for a wide range of emerging applications. Although multimodal sensing systems are essential for capturing complex and dynamic human activities in real-world settings, they bring several new challenges including limited labeled multimodal data. In this paper, we propose Cosmo, a new system for contrastive fusion learning with small data in multimodal HAR applications. Cosmo features a novel two-stage training strategy that leverages both unlabeled data on the cloud and limited labeled data on the edge. By integrating novel fusion-based contrastive learning and quality-guided attention mechanisms, Cosmo can effectively extract both consistent and complementary information across different modalities for efficient fusion. Our evaluation on a cloud-edge testbed using two public datasets and a new multimodal HAR dataset shows that Cosmo delivers significant improvement over state-of-the-art baselines in both recognition accuracy and convergence delay. Xiaomin Ouyang, Xian Shuai, Ivy Wang Shi, Guoliang Xing, Jianwei Huang 0001 |
MobiCom | 1 |
| 2022 | HiToF: a ToF camera system for capturing high-resolution texturesabstractWe present a demonstration of an enhanced Time-of-Flight (ToF) depth system named HiToF, which can expose high-resolution textures from captured depth maps. By design, a ToF camera can easily capture the depth maps of a scene while largely omitting the corresponding texture information, which is often critical for the performance of many depth applications. HiToF is developed to address this issue by generating enhanced depth maps with high-resolution textures. The key idea is to manipulate the phase components used in the measurement of time-of-flight for the received IR light. In this demo, we showcase our implementation using off-the-shelf ToF cameras and engage audience with an interactive experience in various scenarios, which illustrates the system's effectiveness in improving the performance of ToF cameras in depth applications. Xiaomin Ouyang, Li Pan 0004, Wenrui Lu, Xiaoming Liu 0002, Guoliang Xing |
MobiCom | 2 |
| 2021 | ClusterFL: a similarity-aware federated learning system for human activity recognitionabstractFederated Learning (FL) has recently received significant interests thanks to its capability of protecting data privacy. However, existing FL paradigms yield unsatisfactory performance for a wide class of human activity recognition (HAR) applications since they are oblivious to the intrinsic relationship between data of different users. We propose ClusterFL, a similarity-aware federated learning system that can provide high model accuracy and low communication overhead for HAR applications. ClusterFL features a novel clustered multi-task federated learning framework that maximizes the training accuracy of multiple learned models while automatically capturing the intrinsic clustering relationship among the data of different nodes. Based on the learned cluster relationship, ClusterFL can efficiently drop out the nodes that converge slower or have little correlation with other nodes in each cluster, significantly speeding up the convergence while maintaining the accuracy performance. We evaluate the performance of ClusterFL on an NVIDIA edge testbed using four new HAR datasets collected from total 145 users. The results show that, ClusterFL outperforms several state-of-the-art FL paradigms in terms of overall accuracy, and save more than 50% communication overhead at the expense of negligible accuracy degradation. Xiaomin Ouyang, Jianwei Huang 0001, Guoliang Xing |
MobiSys | 1 |
| 2021 | FedDL: Federated Learning via Dynamic Layer Sharing for Human Activity RecognitionabstractDeep learning has been increasingly applied to improve human activity recognition (HAR) accuracy and reduce the human efforts of handcrafted feature extractions. Federated Learning (FL) is an emerging learning paradigm that enables the collaborative learning of a global model without exposing users' raw data. However, existing FL approaches yield unsatisfactory HAR performance as they fail to dynamically aggregate models according to the statistical diversity of users' data. In this paper, we propose FedDL, a novel federated learning system for HAR that can capture the underlying user relationships and apply them to learn personalized models for different users dynamically. Specifically, we design a dynamic layer sharing scheme that learns the similarity among users' model weights to form the sharing structure and merges models accordingly in an iterative, bottom-up layer-wise manner. FedDL merges local models based on the dynamic sharing scheme, significantly speeding up the convergence while maintaining high accuracy. We have implemented FedDL and evaluated using a new data set we collected using LiDAR and four public real-world datasets involving 178 users in total. The results show that FedDL outperforms several state-of-the-art FL paradigms in terms of model accuracy (by more than 15%), converging rate (by more than 70%), and communication overhead (about 30% reduction). Moreover, the testing results on the datasets of different scales show that FedDL has high scalability and hence can be deployed for large-scale real-world applications. Linlin Tu, Xiaomin Ouyang, Guoliang Xing |
SenSys | 2 |
| 2021 | UltraDepth: Exposing High-Resolution Texture from Depth CamerasabstractTime-of-flight (ToF) depth cameras have been increasingly adopted in various real-world applications, e.g., used with RGB cameras for advanced computer vision tasks like 3-D mapping or deployed alone in privacy-sensitive applications such as sleep monitoring. In this paper, we propose UltraDepth, the first system that can expose high-resolution texture from depth maps captured by off-the-shelf ToF cameras, simply by introducing a distorting IR source. The exposed texture information can significantly augment depth-based applications. Moreover, such a capability can be used to launch privacy attacks, which poses a major concern due to the prominence of ToF cameras. To design UltraDepth, we present an in-depth analysis on the impact of the distorting IR light on the distance measurement. We further show that, the reflection properties (reflectivity and incidence angle) of the objects will be encoded in the distorted depth map and hence can be leveraged to reveal texture of objects in UltraDepth. We then propose two practical implementations of UltraDepth, i.e., reflection-based and external IR-based implementations. Our extensive real-world experiments show that, the depth maps output by UltraDepth achieve 89.06%, 99.33%, 81.25% mean accuracy in object detection, face recognition and character recognition, respectively, which offers over 10x improvement over the ordinary depth maps and even approaches the performance of RGB and IR images in a number of scenarios. The findings of this work provide key insights for new research on depth-related computer vision and security of depth sensing devices. Xiaomin Ouyang, Xiaoming Liu 0002, Guoliang Xing |
SenSys | 2 |