EDBT 2026 Demo / reviewers in the wild / expert
Mi Zhang 0002
dblp:84/2519-2
· DBLP profile ↗
58ranked-venue papers
4as first author
35since 2021 · last 2026
0000-0001-7002-6757ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 30 · 20 since 2021Artificial intelligence and machine learning · 16 · 1 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 5 · 2 first-authorSystems, architecture and hardware · 2Applied, interdisciplinary, general and emerging computing · 2 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | RespiraMFM: A Multimodal Foundation Model with Contrastive Audio-Language Alignment for Respiratory Disease IdentificationabstractRespiratory diseases remain a leading cause of global mortality, where timely and accurate diagnosis is critical to improving patient outcomes and reducing healthcare burdens.While prior work has explored audio-based models for respiratory disease detection, such unimodal approaches often suffer from limited generalizability and diagnostic precision.In this paper, we propose RespiraMFM, a Multimodal Foundation Model that integrates respiratory sounds with patient medical history and symptoms to enhance diagnostic accuracy and disease detection capabilities.We introduce an effective contrastive alignment strategy for audio-text multimodal integration, allowing the model to learn better cross-modal representations between respiratory sounds and corresponding textual clinical information.We evaluate RespiraMFM across five major respiratory diseases using seven real-world datasets in both supervised fine-tuning and zero-shot settings, achieving a 9.15% improvement in AU-ROC on supervised tasks and a 20.98% gain on zero-shot tasks over existing baselines.These findings underscore the potential of our framework to advance early diagnosis and improve clinical decision-making in respiratory disease management. Shakhrul Iman Siam, Tiantian Feng, Shri Narayanan, Mi Zhang 0002 |
ACL (1) | 5 |
| 2026 | MemoLens: Empowering Augmented Reality Glasses with Super MemoryabstractSpatial computing empowered by augmented reality (AR) glasses emerges as a new computing paradigm. One of its most transformative capabilities is super memory - the super power to accurately recall objects and people an individual has been paying attention to or interacting with in the physical world. In this work, we propose MemoLensthat empowers AR glasses with super memory. Achieving such capability is not trivial: it requires storing vast amounts of visual data as compact digital memory and retrieving relevant information from the memory efficiently when being prompted. MemoLens achieves this via two key innovations. First, MemoLens leverages eye gaze information captured by the AR glasses to detect what the individual is paying attention to, and introduces a spatio-temporal token compression technique to generate highly compact visual memory. Second, MemoLens introduces a hierarchical search scheme that enables efficient retrieval of relevant memory based on user prompts. We implemented MemoLens using Meta's Aria AR glasses, and evaluated its performance on more than 100 hours of egocentric videos collected in real-world settings. Our results show that MemoLens is able to achieve accurate memory retrieval in real time. Given its promising performance, we believe MemoLens represents a significant step towards realizing always-on super memory in next-generation AR glasses. The project's homepage is https://aiot-mlsys-lab.github.io/memolens.github.io/. Samiul Alam, Shakhrul Iman Siam, Mi Zhang 0002 |
MobiSys | 3 |
| 2026 | GeoFL: A Framework for Efficient Geo-Distributed Cross-Device Federated LearningabstractIn this paper, GeoFL develops a hierarchical federated learning (FL) framework to address the unique challenges in large-scale geo-distributed scenarios. The key idea is to deploy multiple aggregators to geo-distributed clients and aggregate the local model and the global model efficiently and effectively. By assigning each aggregator as a relay layer, GeoFL can elaborately aggregate the geo-distributed clients and systematically determine when to upload the model to the central server based on bandwidth to efficiently update the global model under inadequate and heterogeneous WAN bandwidth constraints. GeoFL designs three key components to optimize the inefficient model aggregation and cope with the non-importance model updates. It further addresses the statistical heterogeneity across geo-distributed aggregators by considering the clients’ graph relationship, delivering an end-to-end clien-taggregator- server architecture for large-scale clients. Compared with existing works, our results on large-scale real-life datasets show that GeoFL speeds up the training process by 1.4×–8× and reduces 6%–80% unnecessary communication rounds between the aggregator and the central server. Maolin Gan, Lanpeng Li, Samiul Alam, Li Liu 0048, Mi Zhang 0002, Huacheng Zeng, Zhichao Cao 0001 |
IEEE Trans. Netw. | 6 |
| 2025 | SVD-LLM: Truncation-aware Singular Value Decomposition for Large Language Model CompressionabstractThe advancements in Large Language Models (LLMs) have been hindered by
their substantial sizes, which necessitates LLM compression methods for practical
deployment. Singular Value Decomposition (SVD) offers a promising solution for
LLM compression. However, state-of-the-art SVD-based LLM compression meth-
ods have two key limitations: truncating smaller singular values may lead to higher
compression loss, and the lack of update on the compressed weights after SVD
truncation. In this work, we propose SVD-LLM, a SVD-based post-training LLM
compression method that addresses the limitations of existing methods. SVD-LLM
incorporates a truncation-aware data whitening technique to ensure a direct map-
ping between singular values and compression loss. Moreover, SVD-LLM adopts
a parameter update with sequential low-rank approximation to compensate for
the accuracy degradation after SVD compression. We evaluate SVD-LLM on 10
datasets and seven models from three different LLM families at three different
scales. Our results demonstrate the superiority of SVD-LLM over state-of-the-arts,
especially at high model compression ratios. Xin Wang 0120, Yu Zheng 0022, Zhongwei Wan, Mi Zhang 0002 |
ICLR | 4 |
| 2025 | D2O: Dynamic Discriminative Operations for Efficient Long-Context Inference of Large Language Models
Zhongwei Wan, Xinjian Wu, Yu Zhang 0133, Yi Xin 0003, Chaofan Tao, Zhihong Zhu 0001, Xin Wang 0120, Longyue Wang, Mi Zhang 0002 |
ICLR | 11 |
| 2025 | GeoFL: A Framework for Efficient Geo-Distributed Cross-Device Federated Learning
Maolin Gan, Lanpeng Li, Samiul Alam, Li Liu 0048, Mi Zhang 0002, Zhichao Cao 0001 |
INFOCOM | 6 |
| 2025 | LoRaSeek: Boosting Denoising Ability in Neural-enhanced LoRa Decoder via Hierarchical Feature ExtractionabstractIn this paper, we propose LoRaSeek, a lightweight and reliable LoRa denoising framework that enhances signal quality and robustness for neural-enhanced LoRa decoding. LoRaSeek integrates a hybrid architecture combining Convolutional Neural Networks (CNNs), Transformers, and a hierarchical U-Net to effectively capture multi-scale, multidimensional features of LoRa chirp signals. To maintain efficiency, we integrate a lightweight Transformer block that supports various LoRa configurations while keeping computational overhead low. Additionally, we incorporate dual attention-based skip connections to preserve chirp signal properties across different scales. Experiments across diverse LoRa configurations show that LoRaSeek achieves 2.04–3.86 dB signal-to-noise ratio (SNR) gains over standard decoding methods and up to 3.03 dB improvement over state-of-the-art neural-enhanced LoRa decoding methods while reducing model storage by up to 7.4× and inference time by up to 1.6×. Yidong Ren, Jialuo Du, Jingkai Lin, Maolin Gan, Shigang Chen, Mi Zhang 0002, Chunyi Peng 0001, Zhichao Cao 0001 |
MobiCom | 7 |
| 2025 | MEDA: Dynamic KV Cache Allocation for Efficient Multimodal Long-Context InferenceabstractZhongwei Wan, Hui Shen, Xin Wang, Che Liu, Zheda Mai, Mi Zhang. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Zhongwei Wan, Hui Shen 0008, Xin Wang 0120, Che Liu 0002, Zheda Mai, Mi Zhang 0002 |
NAACL (Long Papers) | 6 |
| 2025 | SVD-LLM V2: Optimizing Singular Value Truncation for Large Language Model CompressionabstractXin Wang, Samiul Alam, Zhongwei Wan, Hui Shen, Mi Zhang. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Xin Wang 0120, Samiul Alam, Zhongwei Wan, Hui Shen 0008, Mi Zhang 0002 |
NAACL (Long Papers) | 5 |
| 2025 | SRPO: Enhancing Multimodal LLM Reasoning via Reflection-Aware Reinforcement LearningabstractMultimodal large language models (MLLMs) have shown promising capabilities in reasoning tasks, yet still struggle significantly with complex problems requiring explicit self-reflection and self-correction, especially compared to their unimodal text-based counterparts. Existing reflection methods are simplistic and struggle to generate meaningful, instructive feedback, as the reasoning ability and knowledge limits of pre-trained models are largely fixed during initial training. To overcome these challenges, we propose \textit{multimodal \textbf{S}elf-\textbf{R}eflection enhanced reasoning with Group Relative \textbf{P}olicy \textbf{O}ptimization} \textbf{SRPO}, a two-stage reflection-aware reinforcement learning (RL) framework explicitly designed to enhance multimodal LLM reasoning. In the first stage, we construct a high-quality, reflection-focused dataset under the guidance of an advanced MLLM, which generates reflections based on initial responses to help the policy model to learn both reasoning and self-reflection. In the second stage, we introduce a novel reward mechanism within the GRPO framework that encourages concise and cognitively meaningful reflection while avoiding redundancy. Extensive experiments across multiple multimodal reasoning benchmarks—including MathVista, MathVision, Mathverse, and MMMU-Pro—using Qwen-2.5-VL-7B and Qwen-2.5-VL-32B demonstrate that SRPO significantly outperforms state-of-the-art models, achieving notable improvements in both reasoning accuracy and reflection quality. Zhongwei Wan, Zhihao Dou, Che Liu 0002, Yu Zhang 0133, Dongfei Cui, Qinjian Zhao, Hui Shen 0008, Yi Xin 0003, Chaofan Tao, Yangfan He, Mi Zhang 0002, Shen Yan 0008 |
NeurIPS | 13 |
| 2025 | Reading Recognition in the WildabstractTo enable egocentric contextual AI in always-on smart glasses, it is crucial to be able to keep a record of the user's interactions with the world, including during reading. In this paper, we introduce a new task of reading recognition to determine when the user is reading. We first introduce the first-of-its-kind large-scale multimodal Reading in the Wild dataset, containing 100 hours of reading and non-reading videos in diverse and realistic scenarios. We then identify three modalities (egocentric RGB, eye gaze, head pose) that can be used to solve the task, and present a flexible transformer model that performs the task using these modalities, either individually or combined. We show that these modalities are relevant and complementary to the task, and investigate how to efficiently and effectively encode each modality. Additionally, we show the usefulness of this dataset towards classifying types of reading, extending current reading understanding studies conducted in constrained settings to larger scale, diversity and realism. Code, model, and data will be public. Charig Yang, Samiul Alam, Shakhrul Iman Siam, Michael J. Proulx, Lambert Mathias, Kiran K. Somasundaram, Luis Pesqueira, James Fort, Sheroze Sheriffdeen, Omkar M. Parkhi, Carl Yuheng Ren, Mi Zhang 0002, Yuning Chai, Richard A. Newcombe, Hyo Jin Kim 0004 |
NeurIPS | 12 |
| 2025 | Artificial Intelligence of Things: A SurveyabstractThe integration of the Internet of Things (IoT) and modern Artificial Intelligence (AI) has given rise to a new paradigm known as the Artificial Intelligence of Things (AIoT). In this survey, we provide a systematic and comprehensive review of AIoT research. We examine AIoT literature related to sensing, computing, and networking & communication, which form the three key components of AIoT. In addition to advancements in these areas, we review domain-specific AIoT systems that are designed for various important application domains. We have also created an accompanying GitHub repository, where we compile the papers included in this survey: https://github.com/AIoT-MLSys-Lab/AIoT-Survey. This repository will be actively maintained and updated with new research as it becomes available. As both IoT and AI become increasingly critical to our society, we believe that AIoT is emerging as an essential research field at the intersection of IoT and modern AI. It is our hope that this survey will serve as a valuable resource for those engaged in AIoT research and act as a catalyst for future explorations to bridge gaps and drive advancements in this exciting field. Shakhrul Iman Siam, Hyunho Ahn, Li Liu 0048, Samiul Alam, Hui Shen 0008, Zhichao Cao 0001, Ness Shroff, Bhaskar Krishnamachari, Mani Srivastava 0001, Mi Zhang 0002 |
ACM Trans. Sens. Networks | 10 |
| 2024 | ETP: Learning Transferable ECG Representations via ECG-Text Pre-TrainingabstractIn the domain of cardiovascular healthcare, the Electrocardiogram (ECG) serves as a critical, non-invasive diagnostic tool. Although recent strides in self-supervised learning (SSL) have been promising for ECG representation learning, these techniques often require annotated samples and struggle with classes not present in the fine-tuning stages. To address these limitations, we introduce ECG-Text Pre-training (ETP), an innovative framework designed to learn cross-modal representations that link ECG signals with textual reports. For the first time, this framework leverages the zero-shot classification task in the ECG domain. ETP employs an ECG encoder along with a pre-trained language model to align ECG signals with their corresponding textual reports. The proposed framework excels in both linear evaluation and zero-shot classification tasks, as demonstrated on the PTB-XL and CPSC2018 datasets, showcasing its ability for robust and generalizable cross-modal ECG feature learning. Che Liu 0002, Zhongwei Wan, Sibo Cheng, Mi Zhang 0002, Rossella Arcucci |
ICASSP | 4 |
| 2024 | Demeter: Reliable Cross-soil LPWAN with Low-cost Signal Polarization AlignmentabstractSoil monitoring plays an essential role in agricultural systems. Rather than deploying sensors' antennas above the ground, burying them in the soil is an attractive way to retain a non-intrusive aboveground space. Low Power Wide-Area Network (LPWAN) has shown its long-distance and low-power features for aboveground Internet-of-Things (IoT) communication, presenting a potential of extending to underground cross-soil communication over a wide area, which however has not been investigated before. The variation of soil conditions brings significant signal polarization misalignment, degrading communication reliability. In this paper, we propose Demeter, a low-cost low-power programmable antenna design to keep reliable cross-soil communication automatically. First, we propose a hardware architecture to enable polarization adjustment on commercial-off-the-shelf (COTS) single-RF-chain LoRa radio. Moreover, we develop a low-power programmable circuit to obtain polarization adjustment. We further design an energy-efficient heuristic calibration algorithm and an adaptive calibration scheduling method to keep signal polarization alignment automatically. We implement Demeter with a customized PCB circuit and COTS devices. Then, we evaluate its performance in various soil types and environmental conditions. The results show that Demeter can achieve up to 11.6 dB SNR gain indoors and 9.94 dB outdoors, 4× horizontal communication distance, at least 20 cm deeper underground deployment, and up to 82% energy consumption reduction per day compared with the standard LoRa. Yidong Ren, Wei Sun 0002, Jialuo Du, Huaili Zeng, Younsuk Dong, Mi Zhang 0002, Shigang Chen, Yunhao Liu 0001, Tianxing Li 0001, Zhichao Cao 0001 |
MobiCom | 6 |
| 2024 | Demeter-Demo: Demonstrating Cross-soil LPWAN with Low-cost Signal Polarization AlignmentabstractLow Power Wide-Area Network (LPWAN) has shown its long-distance and low-power features for aboveground Internet-of-Things (IoT) communication, presenting a potential to extend to underground cross-soil communication over a wide area, which has not been investigated before. The variation of soil conditions brings significant signal polarization misalignment, degrading communication reliability. We propose Demeter, a low-cost, low-power programmable antenna design to keep reliable cross-soil communication automatically. First, we propose a hardware architecture to enable polarization adjustment on commercial-off-the-shelf (COTS) single-RF-chain LoRa radio. Moreover, we develop a low-power programmable circuit to adjust polarization. We further design an energy-efficient heuristic calibration algorithm to keep signal polarization alignment automatically. We demonstrate Demeter in indoor environments. The antenna is buried in a plastic container filled with gardening soil to simulate the node underground. Meanwhile, we use the COTS LoRa gateway as a receiver to show the RSSI and SNR variations. Yidong Ren, Younsuk Dong, Shigang Chen, Mi Zhang 0002, Jiliang Tang, Zhichao Cao 0001 |
MobiCom | 5 |
| 2024 | ChirpTransformer: Versatile LoRa Encoding for Low-power Wide-area IoTabstractThis paper introduces ChirpTransformer, a versatile LoRa encoding framework that harnesses broad chirp features to dynamically modulate data, enhancing network coverage, throughput, and energy efficiency. Unlike the standard LoRa encoder that offers only single configurable chirp feature, our framework introduces four distinct chirp features, expanding the spectrum of methods available for data modulation. To implement these features on commercial off-the-shelf (COTS) LoRa nodes, we utilize a combination of a software design and a hardware interrupt. ChirpTransformer serves as the foundation for optimizing encoding and decoding in three specific case studies: weak signal decoding for extended network coverage, concurrent transmission for heightened network throughput, and data rate adaptation for improved network energy efficiency. Each case study involves the development of an end-to-end system to comprehensively evaluate its performance. The evaluation results demonstrate remarkable enhancements compared to the standard LoRa. Specifically, ChirpTransformer achieves a 2.38 × increase in network coverage, a 3.14 × boost in network throughput, and a 3.93 × of battery lifetime. Chenning Li, Yidong Ren, Shuai Tong, Shakhrul Iman Siam, Mi Zhang 0002, Jiliang Wang, Yunhao Liu 0001, Zhichao Cao 0001 |
MobiSys | 5 |
| 2024 | WiCGesture: Meta-Motion-Based Continuous Gesture Recognition With Wi-FiabstractRecent advancements in Wi-Fi-based sensing technologies have enabled effective hand gesture recognition. However, most studies focus on single gesture recognition and fail to recognize naturally performed continuous gestures without pauses in transitions. The main challenges include diverse and uncertain transitions in continuous gesture recognition, making it difficult to segment and identify gestures from a stream of continuous hand movements. In this paper, we introduce a new method to recognize continuously performed gestures from a set of predefined gestures (e.g., digits) without requiring a pause in transitions. Instead of segmenting gestures at the gesture-transition level, we segment the stream into basic fractions that depict exclusive moving patterns of gestures. We propose a novel feature called meta motion, which geometrically characterizes different basic hand movements. Leveraging this feature, we use a back-tracking searching-based algorithm to identify gestures from the sequence of meta motions. Based on this approach, we develop a prototype system, WiCGesture, on commodity Wi-Fi devices. WiCGesture is the first system engaging in continuous gesture recognition using Wi-Fi signals. Evaluation results show that WiCGesture effectively recognizes continuous gestures from two gesture sets, significantly outperforming state-of-the-art methods. Ruiyang Gao, Jinyi Liu 0001, Shuyu Dai, Mi Zhang 0002, Leye Wang, Daqing Zhang 0001 |
IEEE Internet Things J. | 5 |
| 2024 | WiVelo: Fine-grained Wi-Fi Walking Velocity EstimationabstractPassive human tracking using Wi-Fi has been researched broadly in the past decade. Besides straightforward anchor point localization, velocity is another vital sign adopted by the existing approaches to infer user trajectory. However, state-of-the-art Wi-Fi velocity estimation relies on Doppler-Frequency-Shift (DFS), which suffers from the inevitable signal noise incurring unbounded velocity errors, further degrading the tracking accuracy. In this article, we present WiVelo, which explores new spatial-temporal signal correlation features observed from different antennas to achieve accurate velocity estimation. First, we use subcarrier shift distribution (SSD) extracted from channel state information (CSI) to define two correlation features for direction and speed estimation, separately. Then, we design a mesh model calculated by the antennas’ locations to enable a fine-grained velocity estimation with bounded direction error. Finally, with the continuously estimated velocity, we develop an end-to-end trajectory recovery algorithm to mitigate velocity outliers with the property of walking velocity continuity. We implement WiVelo on commodity Wi-Fi hardware and extensively evaluate its tracking accuracy in various environments. The experimental results show our median and 90-percentile tracking errors are 0.47 m and 1.06 m, which are half and a quarter of state-of-the-art. The datasets and source codes are published through Github ( https://github.com/research-source/code ). Zhichao Cao 0001, Chenning Li, Li Liu 0048, Mi Zhang 0002 |
ACM Trans. Sens. Networks | 4 |
| 2023 | FedAudio: A Federated Learning Benchmark for Audio TasksabstractFederated learning (FL) has gained substantial attention in recent years due to data privacy concerns related to the pervasiveness of consumer devices that continuously collect data from users. While a number of FL benchmarks have been developed to facilitate FL research, none of them include audio data and audio-related tasks. In this paper, we fill this critical gap by introducing a new FL benchmark for audio tasks which we refer to as FedAudio. FedAudio includes four representative and commonly used audio datasets from three important audio tasks that are well aligned with FL use cases. In particular, a unique contribution of FedAudio is the introduction of data noises and label errors to the datasets to emulate challenges when deploying FL systems in real-world settings. FedAudio also includes the benchmark results of the datasets and a PyTorch library with the objective of facilitating researchers to fairly compare their algorithms. We hope FedAudio could act as a catalyst to inspire new FL research for audio tasks and thus benefit the acoustic and speech research community. The datasets and benchmark results can be accessed at https://github.com/zhang-tuo-pdf/FedAudio. Tiantian Feng, Samiul Alam, Sunwoo Lee 0001, Mi Zhang 0002, Shri Narayanan, Amir Salman Avestimehr |
ICASSP | 5 |
| 2023 | FedMultimodal: A Benchmark for Multimodal Federated LearningabstractOver the past few years, Federated Learning (FL) has become an emerging machine learning technique to tackle data privacy challenges through collaborative training. In the Federated Learning algorithm, the clients submit a locally trained model, and the server aggregates these parameters until convergence. Despite significant efforts that have been made to FL in fields like computer vision, audio, and natural language processing, the FL applications utilizing multimodal data streams remain largely unexplored. It is known that multimodal learning has broad real-world applications in emotion recognition, healthcare, multimedia, and social media, while user privacy persists as a critical concern. Specifically, there are no existing FL benchmarks targeting multimodal applications or related tasks. In order to facilitate the research in multimodal FL, we introduce FedMultimodal, the first FL benchmark for multimodal learning covering five representative multimodal applications from ten commonly used datasets with a total of eight unique modalities. FedMultimodal offers a systematic FL pipeline, enabling end-to-end modeling framework ranging from data partition and feature extraction to FL benchmark algorithms and model evaluation. Unlike existing FL benchmarks, FedMultimodal provides a standardized approach to assess the robustness of FL against three common data corruptions in real-life multimodal applications: missing modalities, missing labels, and erroneous labels. We hope that FedMultimodal can accelerate numerous future research directions, including designing multimodal FL algorithms toward extreme data heterogeneity, robustness multimodal FL, and efficient multimodal FL. The datasets and benchmark results can be accessed at: https://github.com/usc-sail/fed-multimodal. Tiantian Feng, Digbalay Bose, Rajat Hebbar, Anil Ramakrishna, Rahul Gupta 0001, Mi Zhang 0002, Amir Salman Avestimehr, Shri Narayanan |
KDD | 7 |
| 2023 | Med-UniC: Unifying Cross-Lingual Medical Vision-Language Pre-Training by Diminishing BiasabstractThe scarcity of data presents a critical obstacle to the efficacy of medical vision-language pre-training (VLP). A potential solution lies in the combination of datasets from various language communities.
Nevertheless, the main challenge stems from the complexity of integrating diverse syntax and semantics, language-specific medical terminology, and culture-specific implicit knowledge. Therefore, one crucial aspect to consider is the presence of community bias caused by different languages.
This paper presents a novel framework named Unifying Cross-Lingual Medical Vision-Language Pre-Training (\textbf{Med-UniC}), designed to integrate multi-modal medical data from the two most prevalent languages, English and Spanish.
Specifically, we propose \textbf{C}ross-lingual \textbf{T}ext Alignment \textbf{R}egularization (\textbf{CTR}) to explicitly unify cross-lingual semantic representations of medical reports originating from diverse language communities.
\textbf{CTR} is optimized through latent language disentanglement, rendering our optimization objective to not depend on negative samples, thereby significantly mitigating the bias from determining positive-negative sample pairs within analogous medical reports. Furthermore, it ensures that the cross-lingual representation is not biased toward any specific language community.
\textbf{Med-UniC} reaches superior performance across 5 medical image tasks and 10 datasets encompassing over 30 diseases, offering a versatile framework for unifying multi-modal medical data within diverse linguistic communities.
The experimental outcomes highlight the presence of community bias in cross-lingual VLP. Reducing this bias enhances the performance not only in vision-language tasks but also in uni-modal visual tasks. Zhongwei Wan, Che Liu 0002, Mi Zhang 0002, Jie Fu 0001, Benyou Wang, Sibo Cheng, Lei Ma 0008, César Quilodrán Casas, Rossella Arcucci |
NeurIPS | 3 |
| 2023 | Client Selection in Federated Learning: Principles, Challenges, and OpportunitiesabstractAs a privacy-preserving paradigm for training machine learning (ML) models, federated learning (FL) has received tremendous attention from both industry and academia. In a typical FL scenario, clients exhibit significant heterogeneity in terms of data distribution and hardware configurations. Thus, randomly sampling clients in each training round may not fully exploit the local updates from heterogeneous clients, resulting in lower model accuracy, slower convergence rate, degraded fairness, etc. To tackle the FL client heterogeneity problem, various client selection algorithms have been developed, showing promising performance improvement. In this article, we systematically present recent advances in the emerging field of FL client selection and its challenges and research opportunities. We hope to facilitate practitioners in choosing the most suitable client selection mechanisms for their applications, as well as inspire researchers and newcomers to better understand this exciting research topic. Huanle Zhang, Mi Zhang 0002, Xin Liu 0002 |
IEEE Internet Things J. | 4 |
| 2023 | Federated Learning Hyperparameter Tuning From a System PerspectiveabstractFederated learning (FL) is a distributed model training paradigm that preserves clients’ data privacy. It has gained tremendous attention from both academia and industry. FL hyper-parameters (e.g., the number of selected clients and the number of training passes) significantly affect the training overhead in terms of computation time, transmission time, computation load, and transmission load. However, the current practice of manually selecting FL hyper-parameters imposes a heavy burden on FL practitioners because applications have different training preferences. In this paper, we propose, an automatic FL hyper-parameter tuning algorithm tailored to applications’ diverse system requirements in FL training. iteratively adjusts FL hyper-parameters during FL training and can be easily integrated into existing FL systems. Through extensive evaluations of for diverse applications and FL aggregation algorithms, we show that is lightweight and effective, achieving 8.48%-26.75% system overhead reduction compared to using fixed FL hyper-parameters. This paper assists FL practitioners in designing high-performance FL training solutions. The source code of is available at. Huanle Zhang, Mi Zhang 0002, Pengfei Hu 0001, Xiuzhen Cheng, Prasant Mohapatra, Xin Liu 0002 |
IEEE Internet Things J. | 3 |
| 2022 | Multiview Transformers for Video RecognitionabstractVideo understanding requires reasoning at multiple spatiotemporal resolutions – from short fine-grained motions to events taking place over longer durations. Although transformer architectures have recently advanced the state-of-the-art, they have not explicitly modelled different spatiotemporal resolutions. To this end, we present Multiview Transformers for Video Recognition (MTV). Our model consists of separate encoders to represent different views of the input video with lateral connections to fuse information across views. We present thorough ablation studies of our model and show that MTV consistently performs better than single-view counterparts in terms of accuracy and computational cost across a range of model sizes. Furthermore, we achieve state-of-the-art results on six standard datasets, and improve even further with large-scale pretraining. Code and checkpoints are available at: https://github.com/google-research/scenic. Shen Yan 0008, Xuehan Xiong, Anurag Arnab, Zhichao Lu, Mi Zhang 0002, Chen Sun 0002, Cordelia Schmid |
CVPR | 5 |
| 2022 | Deep AutoAugment
Yu Zheng 0022, Shen Yan 0008, Mi Zhang 0002 |
ICLR | 4 |
| 2022 | PyramidFL: a fine-grained client selection framework for efficient federated learningabstractFederated learning (FL) is an emerging distributed machine learning (ML) paradigm with enhanced privacy, aiming to achieve a "good" ML model for as many as participants while consuming as little as wall clock time. By executing across thousands or even millions of clients, FL demonstrates heterogeneous statistical characteristics and system divergence widely across participants, making its training suffer when adopting the traditional ML paradigm. The root cause of the training efficiency degradation is the random client selection criteria. Although existing FL paradigms propose several optimization schemes for client selection, they are still coarse-grained due to their under-exploitation on the clients' data and system heterogeneity, yielding sub-optimal performance for a variety of FL applications. In this paper, we propose PyramidFL1 to speed up the FL training while achieving a higher final model performance (i.e., time-to-accuracy). The core of PyramidFL is a fine-grained client selection, in which PyramidFL does not only focus on the divergence of those selected participants and non-selected ones for client selection but also fully exploits the data and system heterogeneity within selected clients to profile their utility more efficiently. Specifically, PyramidFL first determines the utility-based client selection from the global (i.e., server) view and then optimizes its utility profiling locally (i.e., client) for further client selection. In this way, we can prioritize the use of those clients with higher statistical and system utility consistently. In comparison with the state-of-the-art (i.e., Oort), our evaluation on the open-source FL benchmark shows that PyramidFL improves the final model accuracy by 3.68% -- 7.33%, with a speedup of 2.71 x -- 13.66X on the wall clock time consumption. Chenning Li, Mi Zhang 0002, Zhichao Cao 0001 |
MobiCom | 3 |
| 2022 | FedRolex: Model-Heterogeneous Federated Learning with Rolling Sub-Model ExtractionabstractMost cross-device federated learning (FL) studies focus on the model-homogeneous setting where the global server model and local client models are identical. However, such constraint not only excludes low-end clients who would otherwise make unique contributions to model training but also restrains clients from training large models due to on-device resource bottlenecks. In this work, we propose FedRolex, a partial training (PT)-based approach that enables model-heterogeneous FL and can train a global server model larger than the largest client model. At its core, FedRolex employs a rolling sub-model extraction scheme that allows different parts of the global server model to be evenly trained, which mitigates the client drift induced by the inconsistency between individual client models and server model architectures. Empirically, we show that FedRolex outperforms state-of-the-art PT-based model-heterogeneous FL methods (e.g. Federated Dropout) and reduces the gap between model-heterogeneous and model-homogeneous FL, especially under the large-model large-dataset regime. In addition, we provide theoretical statistical analysis on its advantage over Federated Dropout. Lastly, we evaluate FedRolex on an emulated real-world device distribution to show that FedRolex can enhance the inclusiveness of FL and boost the performance of low-end devices that would otherwise not benefit from FL. Our code is available at: https://github.com/AIoT-MLSys-Lab/FedRolex. Samiul Alam, Ming Yan 0006, Mi Zhang 0002 |
NeurIPS | 4 |
| 2022 | WiVelo: Fine-grained Walking Velocity Estimation for Wi-Fi Passive TrackingabstractPassive human tracking via Wi-Fi has been re-searched broadly in the past decade. Besides straight-forward anchor point localization, velocity is another vital sign adopted by the existing approaches to infer user trajectory. However, state-of-the-art Wi-Fi velocity estimation relies on Doppler-Frequency-Shift (DFS) which suffers from the inevitable signal noise incurring unbounded velocity errors, further degrading the tracking accuracy. In this paper, we present WiVelo11Code&datasets are available at https://github.com/liecn/WiVelo_SECON22 that explores new spatial-temporal signal correlation features observed from different antennas to achieve accurate velocity estimation. First, we use sub carrier shift distribution (SSD) extracted from channel state information (CSI) to define two correlation features for direction and speed estimation, separately. Then, we design a mesh model calculated by the antennas' locations to enable a fine-grained velocity estimation with bounded direction error. Finally, with the continuously estimated velocity, we develop an end-to-end trajectory recovery algorithm to mitigate velocity outliers with the property of walking velocity continuity. We implement WiVelo on commodity Wi-Fi hardware and extensively evaluate its tracking accuracy in various environments. The experimental results show our median and 90% tracking errors are 0.47 m and 1.06 m, which are half and a quarter of state-of-the-arts. Chenning Li, Li Liu 0048, Zhichao Cao 0001, Mi Zhang 0002 |
SECON | 4 |
| 2022 | FedSEA: A Semi-Asynchronous Federated Learning Framework for Extremely Heterogeneous DevicesabstractFederated learning (FL) has attracted increasing attention as a promising technique to drive a vast number of edge devices with artificial intelligence. However, it is very challenging to guarantee the efficiency of a FL system in practice due to the heterogeneous computation resources on different devices. To improve the efficiency of FL systems in the real world, asynchronous FL (AFL) and semi-asynchronous FL (SAFL) methods are proposed such that the server does not need to wait for stragglers. However, existing AFL and SAFL systems suffer from poor accuracy and low efficiency in realistic settings where the data is non-IID distributed across devices and the on-device resources are extremely heterogeneous. In this work, we propose FedSEA - a semi-asynchronous FL framework for extremely heterogeneous devices. We theoretically disclose that the unbalanced aggregation frequency is a root cause of accuracy drop in SAFL. Based on this analysis, we design a training configuration scheduler to balance the aggregation frequency of devices such that the accuracy can be improved. To improve the efficiency of the system in realistic settings where the devices have dynamic on-device resource availability, we design a scheduler that can efficiently predict the arriving time of local updates from devices and adjust the synchronization time point according to the devices' predicted arriving time. We also consider the extremely heterogeneous settings where there exist extremely lagging devices that take hundreds of times as long as the training time of the other devices. In the real world, there might be even some extreme stragglers which are not capable of training the global model. To enable these devices to join in training without impairing the systematic efficiency, Fed-SEA enables these extreme stragglers to conduct local training on much smaller models. Our experiments show that compared with status quo approaches, FedSEA improves the inference accuracy by 44.34% and reduces the systematic time cost and local training time cost by 87.02× and 792.9×. FedSEA also reduces the energy consumption of the devices with extremely limited resources by 752.9×. Jingwei Sun 0002, Ang Li 0005, Lin Duan, Samiul Alam, Xuliang Deng, Xin Guo 0008, Haiming Wang 0002, Maria Gorlatova, Mi Zhang 0002, Hai Li 0001, Yiran Chen 0001 |
SenSys | 9 |
| 2021 | CATE: Computation-aware Neural Architecture Encoding with TransformersabstractRecent works (White et al., 2020a; Yan et al., 2020) demonstrate the importance of architecture encodings in Neural Architecture Search (NAS). These encodings encode either structure or computation information of the neural architectures. Compared to structure-aware encodings, computation-aware encodings map architectures with similar accuracies to the same region, which improves the downstream architecture search performance (Zhang et al., 2019; White et al., 2020a). In this work, we introduce a Computation-Aware Transformer-based Encoding method called CATE. Different from existing computation-aware encodings based on fixed transformation (e.g. path encoding), CATE employs a pairwise pre-training scheme to learn computation-aware encodings using Transformers with cross-attention. Such learned encodings contain dense and contextualized computation information of neural architectures. We compare CATE with eleven encodings under three major encoding-dependent NAS subroutines in both small and large search spaces. Our experiments show that CATE is beneficial to the downstream search, especially in the large search space. Moreover, the outside search space experiment demonstrates its superior generalization ability beyond the search space on which it was trained. Our code is available at: https://github.com/MSU-MLSys-Lab/CATE. Shen Yan 0008, Kaiqiang Song, Fei Liu 0004, Mi Zhang 0002 |
ICML | 4 |
| 2021 | DeepLoRa: Learning Accurate Path Loss Model for Long Distance Links in LPWANabstractLoRa (Long Range) is an emerging wireless technology that enables long-distance communication and keeps low power consumption. Therefore, LoRa plays a more and more important role in Low-Power Wide-Area Networks (LPWANs), which easily extend many large-scale Internet of Things (IoT) applications in diverse scenarios (e.g., industry, agriculture, city). In lots of environments where various types of land-covers usually exist, it is challenging to precisely predict a LoRa link's path loss. As a result, how to deploy LoRa gateways to ensure reliable coverage and develop precise fingerprint-based localization becomes a difficult issue in practice. In this paper, we propose DeepLoRa, a deep learning-based approach to accurately estimate the path loss of long-distance links in complex environments. Specifically, DeepLoRa relies on remote sensing to automatically recognize land-cover types along a LoRa link. Then, DeepLoRa utilizes Bi-LSTM (Bidirectional Long Short Term Memory) to develop a land-cover aware path loss model. We implement DeepLoRa and use the data gathered from a real LoRaWAN deployment on campus to evaluate its performance extensively in terms of estimation accuracy and model transferability. The results show that DeepLoRa reduces the estimation error to less than 4 dB, which is 2× smaller than state-of-the-art models. Li Liu 0048, Yuguang Yao, Zhichao Cao 0001, Mi Zhang 0002 |
INFOCOM | 4 |
| 2021 | FedMask: Joint Computation and Communication-Efficient Personalized Federated Learning via Heterogeneous MaskingabstractRecent advancements in deep neural networks (DNN) enabled various mobile deep learning applications. However, it is technically challenging to locally train a DNN model due to limited data on devices like mobile phones. Federated learning (FL) is a distributed machine learning paradigm which allows for model training on decentralized data residing on devices without breaching data privacy. Hence, FL becomes a natural choice for deploying on-device deep learning applications. However, the data residing across devices is intrinsically statistically heterogeneous (i.e., non-IID data distribution) and mobile devices usually have limited communication bandwidth to transfer local updates. Such statistical heterogeneity and communication bandwidth limit are two major bottlenecks that hinder applying FL in practice. In addition, considering mobile devices usually have limited computational resources, improving computation efficiency of training and running DNNs is critical to developing on-device deep learning applications. In this paper, we present FedMask - a communication and computation efficient FL framework. By applying FedMask, each device can learn a personalized and structured sparse DNN, which can run efficiently on devices. To achieve this, each device learns a sparse binary mask (i.e., 1 bit per network parameter) while keeping the parameters of each local model unchanged; only these binary masks will be communicated between the server and the devices. Instead of learning a shared global model in classic FL, each device obtains a personalized and structured sparse model that is composed by applying the learned binary mask to the fixed parameters of the local model. Our experiments show that compared with status quo approaches, FedMask improves the inference accuracy by 28.47% and reduces the communication cost and the computation cost by 34.48X and 2.44X. FedMask also achieves 1.56X inference speedup and reduces the energy consumption by 1.78X. Ang Li 0005, Jingwei Sun 0002, Mi Zhang 0002, Hai Li 0001, Yiran Chen 0001 |
SenSys | 4 |
| 2021 | NELoRa: Towards Ultra-low SNR LoRa Communication with Neural-enhanced DemodulationabstractLow-Power Wide-Area Networks (LPWANs) are an emerging Internet-of-Things (IoT) paradigm marked by low-power and long-distance communication. Among them, LoRa is widely deployed for its unique characteristics and open-source technology. By adopting the Chirp Spread Spectrum (CSS) modulation, LoRa enables low signal-to-noise ratio (SNR) communication. However, the standard demodulation method does not fully exploit the properties of chirp signals, thus yields a sub-optimal SNR threshold under which the decoding fails. Consequently, the communication range and energy consumption have to be compromised for robust transmission. This paper presents NELoRa, a neural-enhanced LoRa demodulation method, exploiting the feature abstraction ability of deep learning to support ultra-low SNR LoRa communication. Taking the spectrogram of both amplitude and phase as input, we first design a mask-enabled Deep Neural Network (DNN) filter that extracts multi-dimension features to capture clean chirp symbols. Second, we develop a spectrogram-based DNN decoder to decode these chirp symbols accurately. Finally, we propose a generic packet demodulation system by incorporating a method that generates high-quality chirp symbols from received signals. We implement and evaluate NELoRa on both indoor and campus-scale outdoor testbeds. The results show that NELoRa achieves 1.84-2.35 dB SNR gains and extends the battery life up to 272% (~0.38-1.51 years) in average for various LoRa configurations. Chenning Li, Hanqing Guo, Shuai Tong, Zhichao Cao 0001, Mi Zhang 0002, Qiben Yan 0001, Li Xiao 0001, Jiliang Wang, Yunhao Liu 0001 |
SenSys | 6 |
| 2021 | Mercury: Efficient On-Device Distributed DNN Training via Stochastic Importance SamplingabstractAs intelligence is moving from data centers to the edges, intelligent edge devices such as smartphones, drones, robots, and smart IoT devices are equipped with the capability to altogether train a deep learning model on the devices from the data collected by themselves. Despite its considerable value, the key bottleneck of making on-device distributed training practically useful in real-world deployments is that they consume a significant amount of training time under wireless networks with constrained bandwidth. To tackle this critical bottleneck, we present Mercury, an importance sampling-based framework that enhances the training efficiency of on-device distributed training without compromising the accuracies of the trained models. The key idea behind the design of Mercury is to focus on samples that provide more important information in each training iteration. In doing this, the training efficiency of each iteration is improved. As such, the total number of iterations can be considerably reduced so as to speed up the overall training process. We implemented Mercury and deployed it on a self-developed testbed. We demonstrate its effectiveness and show that Mercury consistently outperforms two status quo frameworks on six commonly used datasets across tasks in image classification, speech recognition, and natural language processing. Ming Yan 0006, Mi Zhang 0002 |
SenSys | 3 |
| 2021 | The Untold Secrets of WiFi-Calling Services: Vulnerabilities, Attacks, and CountermeasuresabstractSince 2016, all of four major U.S. operators have rolled out Wi-Fi calling services. They enable mobile users to place cellular calls over Wi-Fi networks based on the 3GPP IMS technology. Compared with conventional cellular voice solutions, the major difference lies in that their traffic traverses untrusted Wi-Fi networks and the Internet. This exposure to insecure networks can cause the Wi-Fi calling users to suffer from security threats. Its security mechanisms are similar to the VoLTE, because both of them are supported by the IMS. They include SIM-based security, 3GPP AKA, IPSec, etc. However, are they sufficient to secure Wi-Fi calling services? Unfortunately, our study yields a negative answer. We conduct the first security study on the operational Wi-Fi calling services in three major U.S. operators networks using commodity devices. We disclose that current Wi-Fi calling security is not bullet-proof and uncover three vulnerabilities. By exploiting the vulnerabilities, we devise two proof-of-concept attacks: telephony harassment or denial of voice service and user privacy leakage; both of them can bypass the existing security defenses. We have confirmed their feasibility using real-world experiments, as well as assessed their potential damages and proposed a solution to address all identified vulnerabilities. Tian Xie 0001, Guan-Hua Tu, Bangjie Yin, Chi-Yu Li 0001, Chunyi Peng 0001, Mi Zhang 0002, Hui Liu 0031, Xiaoming Liu 0002 |
IEEE Trans. Mob. Comput. | 6 |
| 2020 | MutualNet: Adaptive ConvNet via Mutual Learning from Network Width and Resolution
Taojiannan Yang, Sijie Zhu, Chen Chen 0001, Shen Yan 0008, Mi Zhang 0002, Andrew R. Willis |
ECCV (1) | 5 |
| 2020 | Deep Subclass Linear Discriminant Analysis For Multimodal Feature Space LearningabstractIn this work, we target a known problem in representation learning that is: beyond coarse classification, how can we better model fine-grained categorization? To address this problem, we introduce Deep Subclass Linear Discriminant Analysis (DeepSDA), which utilizes intra-class variation and inter-class similarity during training. We could achieve multimodal classification by maximizing the ratio of between-subclass scatter matrix and within-subclass scatter matrix. We maximize the eigenvalues along the discriminative eignevector directions. Hence the deep neural network is able to learn more discriminative representation space and thus has higher class separation in the linearly separable latent space. We show that DeepSDA leads to significant improvements on diverse fine-grained categorization and attribute learning benchmarks. Abin Jose, Shen Yan 0008, Mi Zhang 0002, Jens-Rainer Ohm |
ICIP | 3 |
| 2020 | FlexDNN: Input-Adaptive On-Device Deep Learning for Efficient Mobile VisionabstractMobile vision systems powered by the recent advancement in Deep Neural Networks (DNNs) are enabling a wide range of on-device video analytics applications. Considering mobile systems are constrained with limited resources, reducing resource demands of DNNs is crucial to realizing the full potential of these applications. In this paper, we present FlexDNN, an input-adaptive DNN-based framework for efficient on-device video analytics. To achieve this, FlexDNN takes the intrinsic dynamics of mobile videos into consideration, and dynamically adapts its model complexity to the difficulty levels of input video frames to achieve computation efficiency. FlexDNN addresses the key drawbacks of existing systems and pushes the state-of-the-art forward. We use FlexDNN to build three representative on-device video analytics applications, and evaluate its performance on both mobile CPU and GPU platforms. Our results show that FlexDNN significantly outperforms status quo approaches in accuracy, average CPU/GPU processing time per frame, frame drop rate, and energy consumption. Biyi Fang, Faen Zhang, Mi Zhang 0002 |
SEC | 5 |
| 2020 | SCYLLA: QoE-aware Continuous Mobile Vision with FPGA-based Dynamic Deep Neural Network ReconfigurationabstractContinuous mobile vision is becoming increasingly important as it finds compelling applications which substantially improve our everyday life. However, meeting the requirements of quality of experience (QoE) diversity, energy efficiency and multi-tenancy simultaneously represents a significant challenge. In this paper, we present SCYLLA, an FPGA-based framework that enables QoE-aware continuous mobile vision with dynamic reconfiguration to effectively address this challenge. SCYLLA pre-generates a pool of FPGA design and DNN models, and dynamically applies the optimal software-hardware configuration to achieve the maximum overall performance on QoE for concurrent tasks. We implement SCYLLA on state-of-the-art FPGA platform and evaluate SCYLLA using drone-based traffic surveillance application on three datasets. Our evaluation shows that SCYLLA provides much better design flexibility and achieves superior QoE trade-offs than status-quo CPU-based solution that existing continuous mobile vision applications are built upon. Shuang Jiang, Zhiyao Ma, Chenren Xu, Mi Zhang 0002, Chen Zhang 0001, Yunxin Liu 0001 |
INFOCOM | 5 |
| 2020 | SecWIR: securing smart home IoT communications via wi-fi routers with embedded intelligenceabstractSmart home Wi-Fi IoT devices are prevalent nowadays and potentially bring significant improvements to daily life. However, they pose an attractive target for adversaries seeking to launch attacks. Since the secure IoT communications are the foundation of secure IoT devices, this study commences by examining the extent to which mainstream security protocols are supported by 40 of the best selling Wi-Fi smart home IoT devices on the Amazon platform. It is shown that 29 of these devices have either no security protocols deployed, or have problematic security protocol implementations. Seemingly, these vulnerabilities can be easily fixed by installing security patches. However, many IoT devices lack the requisite software/hardware resources to do so. To address this problem, the present study proposes a SecWIR (Secure Wi-Fi IoT communication Router) framework designed for implementation on top of the users' existing home Wi-Fi routers to provide IoT devices with a secure IoT communication capability. However, it is way challenging for SecWIR to function effectively on all home Wi-Fi routers since some routers are resource-constrained. Thus, several novel techniques for resolving this implementation issue are additionally proposed. The experimental results show that SecWIR performs well on a variety of commercial off-the-shelf (COTS) Wi-Fi routers at the expense of only a small reduction in the non-IoT data service throughput (less than 8%), and small increases in the CPU usage (4.5%~7%), RAM usage (1.9 MB~2.2 MB), and the IoT device access delay (24 ms~154 ms) while securing 250 IoT devices. Guan-Hua Tu, Chi-Yu Li 0001, Tian Xie 0001, Mi Zhang 0002 |
MobiSys | 5 |
| 2020 | Does Unsupervised Architecture Representation Learning Help Neural Architecture Search?abstractExisting Neural Architecture Search (NAS) methods either encode neural architectures using discrete encodings that do not scale well, or adopt supervised learning-based methods to jointly learn architecture representations and optimize architecture search on such representations which incurs search bias. Despite the widespread use, architecture representations learned in NAS are still poorly understood. We observe that the structural properties of neural architectures are hard to preserve in the latent space if architecture representation learning and search are coupled, resulting in less effective search performance. In this work, we find empirically that pre-training architecture representations using only neural architectures without their accuracies as labels improves the downstream architecture search efficiency. To explain this finding, we visualize how unsupervised architecture representation learning better encourages neural architectures with similar connections and operators to cluster together. This helps map neural architectures with similar performance to the same regions in the latent space and makes the transition of architectures in the latent space relatively smooth, which considerably benefits diverse downstream search strategies. Shen Yan 0008, Yu Zheng 0022, Mi Zhang 0002 |
NeurIPS | 5 |
| 2020 | Wi-fi see it all: generative adversarial network-augmented versatile wi-fi imagingabstractWi-Fi imaging has attracted significant interests due to the ubiquitous availability of Wi-Fi devices today. In this paper, we present Wi-Fi See It All (WiSIA), a versatile Wi-Fi imaging system built upon commercial off-the-shelf (COTS) Wi-Fi devices, which is able to simultaneously detect objects and humans, segment their boundaries, and identify them within the image plane. To achieve this, WiSIA utilizes three techniques. First, instead of constructing the image plane at the receiver side using a high-cost antenna array and complex parameter estimation, WiSIA pushes the image plane to the object side with two pairs of transceivers and 2D-IFFT. Second, WiSIA extracts the specific physical signature of the signals reflected from multiple objects to segment their boundaries. Third, WiSIA incorporates a cGAN (conditional Generative Adversarial Network) to enhance the boundary of different objects. We have implemented WiSIA using COTS Wi-Fi devices and evaluated it using a rich set of experiments. Our results demonstrate the efficacy of WiSIA. It outperforms the state-of-the-art vision-based method in dark and occlusion scenarios, demonstrating its superiority in such challenge scenarios. Chenning Li, Yuguang Yao, Zhichao Cao 0001, Mi Zhang 0002, Yunhao Liu 0001 |
SenSys | 5 |
| 2020 | Distream: scaling live video analytics with workload-adaptive distributed edge intelligenceabstractVideo cameras have been deployed at scale today. Driven by the breakthrough in deep learning (DL), organizations that have deployed these cameras start to use DL-based techniques for live video analytics. Although existing systems aim to optimize live video analytics from a variety of perspectives, they are agnostic to the workload dynamics in real-world deployments. In this work, we present Distream, a distributed live video analytics system based on the smart camera-edge cluster architecture, that is able to adapt to the workload dynamics to achieve low-latency, high-throughput, and scalable live video analytics. The key behind the design of Distream is to adaptively balance the workloads across smart cameras and partition the workloads between cameras and the edge cluster. In doing so, Distream is able to fully utilize the compute resources at both ends to achieve optimized system performance. We evaluated Distream with 500 hours of distributed video streams from two real-world video datasets with a testbed that consists of 24 cameras and a 4-GPU edge cluster. Our results show that Distream consistently outperforms the status quo in terms of throughput, latency, and latency service level objective (SLO) miss rate. Biyi Fang, Haichen Shen, Mi Zhang 0002 |
SenSys | 4 |
| 2018 | NestDNN: Resource-Aware Multi-Tenant On-Device Deep Learning for Continuous Mobile VisionabstractMobile vision systems such as smartphones, drones, and augmented-reality headsets are revolutionizing our lives. These systems usually run multiple applications concurrently and their available resources at runtime are dynamic due to events such as starting new applications, closing existing applications, and application priority changes. In this paper, we present NestDNN, a framework that takes the dynamics of runtime resources into account to enable resource-aware multi-tenant on-device deep learning for mobile vision systems. NestDNN enables each deep learning model to offer flexible resource-accuracy trade-offs. At runtime, it dynamically selects the optimal resource-accuracy trade-off for each deep learning model to fit the model's resource demand to the system's available runtime resources. In doing so, NestDNN efficiently utilizes the limited resources in mobile vision systems to jointly maximize the performance of all the concurrently running applications. Our experiments show that compared to the resource-agnostic status quo approach, NestDNN achieves as much as 4.2% increase in inference accuracy, 2.0× increase in video frame processing rate and 1.7× reduction on energy consumption. Biyi Fang, Mi Zhang 0002 |
MobiCom | 3 |
| 2017 | MobileDeepPill: A Small-Footprint Mobile Deep Learning System for Recognizing Unconstrained Pill ImagesabstractCorrect identification of prescription pills based on their visual appearance is a key step required to assure patient safety and facilitate more effective patient care. With the availability of high-quality cameras and computational power on smartphones, it is possible and helpful to identify unknown prescription pills using smartphones. Towards this goal, in 2016, the U.S. National Library of Medicine (NLM) of the National Institutes of Health (NIH) announced a nationwide competition, calling for the creation of a mobile vision system that can recognize pills automatically from a mobile phone picture under unconstrained real-world settings. In this paper, we present the design and evaluation of such mobile pill image recognition system called MobileDeepPill. The development of MobileDeepPill involves three key innovations: a triplet loss function which attains invariances to real-world noisiness that deteriorates the quality of pill images taken by mobile phones; a multi-CNNs model that collectively captures the shape, color and imprints characteristics of the pills; and a Knowledge Distillation-based deep model compression framework that significantly reduces the size of the multi-CNNs model without deteriorating its recognition performance. Our deep learning-based pill image recognition algorithm wins the First Prize (champion) of the NIH NLM Pill Image Recognition Challenge. Given its promising performance, we believe MobileDeepPill helps NIH tackle a critical problem with significant societal impact and will benefit millions of healthcare personnel and the general public. Mi Zhang 0002 |
MobiSys | 3 |
| 2017 | DeepASL: Enabling Ubiquitous and Non-Intrusive Word and Sentence-Level Sign Language TranslationabstractThere is an undeniable communication barrier between deaf people and people with normal hearing ability. Although innovations in sign language translation technology aim to tear down this communication barrier, the majority of existing sign language translation systems are either intrusive or constrained by resolution or ambient lighting conditions. Moreover, these existing systems can only perform single-sign ASL translation rather than sentence-level translation, making them much less useful in daily-life communication scenarios. In this work, we fill this critical gap by presenting DeepASL, a transformative deep learning-based sign language translation technology that enables ubiquitous and non-intrusive American Sign Language (ASL) translation at both word and sentence levels. DeepASL uses infrared light as its sensing mechanism to non-intrusively capture the ASL signs. It incorporates a novel hierarchical bidirectional deep recurrent neural network (HB-RNN) and a probabilistic framework based on Connectionist Temporal Classification (CTC) for word-level and sentence-level ASL translation respectively. To evaluate its performance, we have collected 7, 306 samples from 11 participants, covering 56 commonly used ASL words and 100 ASL sentences. DeepASL achieves an average 94.5% word-level translation accuracy and an average 8.2% word error rate on translating unseen ASL sentences. Given its promising performance, we believe DeepASL represents a significant step towards breaking the communication barrier between deaf people and hearing majority, and thus has the significant potential to fundamentally change deaf people's lives. Biyi Fang, Jillian Co, Mi Zhang 0002 |
SenSys | 3 |
| 2016 | AirSense: an intelligent home-based sensing system for indoor air quality analyticsabstractIn the U.S., people spend approximately 90 percent of their time indoors. Unfortunately, indoor air quality (IAQ) may be two to five times worse than the air outdoors, and is often overlooked. Existing IAQ monitoring technologies focus on IAQ measurements and visualization. However, the lack of information about the pollution sources as well as the seriousness of the pollution makes people feel powerless and frustrated, resulting in the ignorance of the polluted air at their homes. In this work, we fill this critical gap by presenting AirSense, an intelligent home-based IAQ sensing system that is able to automatically detect pollution events, identify pollution sources, estimate personal exposure to indoor air pollution, and provide actionable suggestions to help people improve IAQ. We have deployed AirSense at five homes to evaluate its performance and investigate how users interact with it. We demonstrate that AirSense can accurately detect pollution events, identify pollution sources, and forecast IAQ information within five minutes in both controlled and real-world settings. We further show the great potential of AirSense in increasing users' awareness of IAQ and helping them better manage IAQ at their homes. Biyi Fang, Qiumin Xu, Taiwoo Park, Mi Zhang 0002 |
UbiComp | 4 |
| 2016 | HeadScan: A Wearable System for Radio-Based Sensing of Head and Mouth-Related ActivitiesabstractThe popularity of wearables continues to rise. However, their functionalities and applications are constrained by the types of sensors that are currently available. Accelerometers and gyroscopes struggle to capture complex user activities. Microphones and image sensors are more powerful but capture privacy sensitive information. Physiological sensors are obtrusive to users since they often require skin contact and must be placed at certain body positions to function. In contrast, radio- based sensing uses wireless radio signals to capture movements of different parts of body caused by human activities and therefore provides a contactless and privacy-preserving approach to detect and monitor human activities. In this paper, we contribute to the search for a new sensing modality for the next generation of wearable devices by exploring the feasibility of radio-based human activity sensing and recognition in the context of wearable setting. We envision radio-based sensing has the potential to fundamentally transform wearables as we currently know them. As the first step to achieve our vision, we have designed and developed HeadScan, a first- of-its-kind wearable for radio-based sensing of a number of human activities that involve head and mouth movements. HeadScan only requires a pair of small antennas placed on the shoulder and collar and one wearable unit worn on the arm or the belt of the user. HeadScan uses the fine-grained CSI measurements extracted from the radio signals and incorporates a radio signal processing pipeline that converts the raw CSI measurements into the targeted human activities. To examine the feasibility and performance of HeadScan, we have collected about 50.5 hours data from seven users. Our wide-range experiments including comparisons to a conventional skin-contact audio-based sensing approach to tracking the same set of head and mouth-related activities highlight the enormous potential of our radio-based sensing approach and provide guidance to future explorations. Biyi Fang, Nicholas D. Lane, Mi Zhang 0002, Fahim Kawsar |
IPSN | 3 |
| 2016 | BodyScan: Enabling Radio-based Sensing on Wearable Devices for Contactless Activity and Vital Sign MonitoringabstractWearable devices are increasingly becoming mainstream consumer products carried by millions of consumers. However, the potential impact of these devices is currently constrained by fundamental limitations of their built-in sensors. In this paper, we introduce radio as a new powerful sensing modality for wearable devices and propose to transform radio into a mobile sensor of human activities and vital signs. We present BodyScan, a wearable system that enables radio to act as a single modality capable of providing whole-body continuous sensing of the user. BodyScan overcomes key limitations of existing wearable devices by providing a contactless and privacy-preserving approach to capturing a rich variety of human activities and vital sign information. Our prototype design of BodyScan is comprised of two components: one worn on the hip and the other worn on the wrist, and is inspired by the increasingly prevalent scenario where a user carries a smartphone while also wearing a wristband/smartwatch. This prototype can support daily usage with one single charge per day. Experimental results show that in controlled settings, BodyScan can recognize a diverse set of human activities while also estimating the user's breathing rate with high accuracy. Even in very challenging real-world settings, BodyScan can still infer activities with an average accuracy above 60% and monitor breathing rate information a reasonable amount of time during each day. Biyi Fang, Nicholas D. Lane, Mi Zhang 0002, Aidan Boran, Fahim Kawsar |
MobiSys | 3 |
| 2015 | MyBehavior: automatic personalized health feedback from user behaviors and preferences using smartphonesabstractMobile sensing systems have made significant advances in tracking human behavior. However, the development of personalized mobile health feedback systems is still in its infancy. This paper introduces MyBehavior, a smartphone application that takes a novel approach to generate deeply personalized health feedback. It combines state-of-the-art behavior tracking with algorithms that are used in recommendation systems. MyBehavior automatically learns a user's physical activity and dietary behavior and strategically suggests changes to those behaviors for a healthier lifestyle. The system uses a sequential decision making algorithm, Multi-armed Bandit, to generate suggestions that maximize calorie loss and are easy for the user to adopt. In addition, the system takes into account user's preferences to encourage adoption using the pareto-frontier algorithm. In a 14-week study, results show statistically significant increases in physical activity and decreases in food calorie when using MyBehavior compared to a control condition. Mashfiqui Rabbi, M. S. Hane Aung, Mi Zhang 0002, Tanzeem Choudhury |
UbiComp | 3 |
| 2015 | DoppleSleep: a contactless unobtrusive sleep sensing system using short-range Doppler radarabstractIn this paper, we present DoppleSleep -- a contactless sleep sensing system that continuously and unobtrusively tracks sleep quality using commercial off-the-shelf radar modules. DoppleSleep provides a single sensor solution to track sleep-related physical and physiological variables including coarse body movements and subtle and fine-grained chest, heart movements due to breathing and heartbeat. By integrating vital signals and body movement sensing, DoppleSleep achieves 89.6% recall with Sleep vs. Wake classification and 80.2% recall with REM vs. Non-REM classification compared to EEG-based sleep sensing. Lastly, it provides several objective sleep quality measurements including sleep onset latency, number of awakenings, and sleep efficiency. The contactless nature of DoppleSleep obviates the need to instrument the user's body with sensors. Lastly, DoppleSleep is implemented on an ARM microcontroller and a smartphone application that are benchmarked in terms of power and resource usage. Tauhidur Rahman, Alexander Travis Adams, Ruth Vinisha, Mi Zhang 0002, Shwetak N. Patel, Julie A. Kientz, Tanzeem Choudhury |
UbiComp | 4 |
| 2015 | Feasibility of B-mode diagnostic ultrasonic energy transfer and telemetry to a cm2 sized deep-tissue implantabstractWhile radio-frequency based remote powering and back-telemetry is popular for many of the surface implants, like neural prosthesis or under-the-skin implanted batteries, it is not suitable for implants located deep inside the tissue. In this paper, we investigate the feasibility of using a commercial off-the-shelf (COTS), diagnostic ultrasound technology for delivering energy to a sub-cm2sized device implanted at depths more than 10cm away from the tissue surface. Using a COTS 3.5MHz ultrasound scanner we show how the B-mode interrogation protocol can be used to deliver energy to an encapsulated PZT transducer and how the B-mode video sequence can be parsed to retrieve the data from the transducer. In this paper we also discuss the limits of energy transfer at different implantation depths and we also discuss the energy requirements at the implant to achieve robust data transfer. Biyi Fang, Mi Zhang 0002, Shantanu Chakrabartty |
ISCAS | 3 |
| 2014 | BodyBeat: a mobile system for sensing non-speech body soundsabstractIn this paper, we propose BodyBeat, a novel mobile sensing system for capturing and recognizing a diverse range of non-speech body sounds in real-life scenarios. Non-speech body sounds, such as sounds of food intake, breath, laughter, and cough contain invaluable information about our dietary behavior, respiratory physiology, and affect. The BodyBeat mobile sensing system consists of a custom-built piezoelectric microphone and a distributed computational framework that utilizes an ARM microcontroller and an Android smartphone. The custom-built microphone is designed to capture subtle body vibrations directly from the body surface without being perturbed by external sounds. The microphone is attached to a 3D printed neckpiece with a suspension mechanism. The ARM embedded system and the Android smartphone process the acoustic signal from the microphone and identify non-speech body sounds. We have extensively evaluated the BodyBeat mobile sensing system. Our results show that BodyBeat outperforms other existing solutions in capturing and recognizing different types of important non-speech body sounds. Tauhidur Rahman, Alexander Travis Adams, Mi Zhang 0002, Erin Cherry, Bobby Zhou, Huaishu Peng, Tanzeem Choudhury |
MobiSys | 3 |
| 2013 | Human Daily Activity Recognition With Sparse Representation Using Wearable SensorsabstractHuman daily activity recognition using mobile personal sensing technology plays a central role in the field of pervasive healthcare. One major challenge lies in the inherent complexity of human body movements and the variety of styles when people perform a certain activity. To tackle this problem, in this paper, we present a novel human activity recognition framework based on recently developed compressed sensing and sparse representation theory using wearable inertial sensors. Our approach represents human activity signals as a sparse linear combination of activity signals from all activity classes in the training set. The class membership of the activity signal is determined by solving a l(1) minimization problem. We experimentally validate the effectiveness of our sparse representation-based approach by recognizing nine most common human daily activities performed by 14 subjects. Our approach achieves a maximum recognition rate of 96.1%, which beats conventional methods based on nearest neighbor, naive Bayes, and support vector machine by as much as 6.7%. Furthermore, we demonstrate that by using random projection, the task of looking for “optimal features” to achieve the best activity recognition performance is less important within our framework. Mi Zhang 0002, Alexander A. Sawchuk |
IEEE J. Biomed. Health Informatics | 1 |
| 2012 | Co-recognition of Human Activity and Sensor Location via Compressed Sensing in Wearable Body Sensor NetworksabstractHuman activity recognition using wearable body sensors is playing a significant role in ubiquitous and mobile computing. One of the issues related to this wearable technology is that the captured activity signals are highly dependent on the location where the sensors are worn on the human body. Existing research work either extracts location information from certain activity signals or takes advantage of the sensor location information as a priori to achieve better activity recognition performance. In this paper, we present a compressed sensing-based approach to co-recognize human activity and sensor location in a single framework. To validate the effectiveness of our approach, we did a pilot study for the task of recognizing 14 human activities and 7 on body-locations. On average, our approach achieves an 87:72% classification accuracy (the mean of precision and recall). Wenyao Xu, Mi Zhang 0002, Alexander A. Sawchuk, Majid Sarrafzadeh |
BSN | 2 |
| 2012 | A preliminary study of sensing appliance usage for human activity recognition using mobile magnetometerabstractHuman activity recognition and human behavior understanding play a central role in the field of ubiquitous computing. In this paper, we propose a novel method using magnetometer embedded in the mobile phone to recognize activities by detecting household appliance usage. The key idea of our approach is that when the mobile phone user performs a certain activity at home, the embedded magnetometer is capable of capturing the changes of the magnetic field strength around the mobile phone caused by the household appliance in operation. Our mobile application uses these changes as magnetic signatures for each of these appliance such that the daily household acitivities associated with these appliance such as cooking can be recognized. Mi Zhang 0002, Alexander A. Sawchuk |
UbiComp | 1 |
| 2012 | USC-HAD: a daily activity dataset for ubiquitous activity recognition using wearable sensorsabstractMany ubiquitous computing applications involve human activity recognition based on wearable sensors. Although this problem has been studied for a decade, there are a limited number of publicly available datasets to use as standard benchmarks to compare the performance of activity models and recognition algorithms. In this paper, we describe the freely available USC human activity dataset (USC-HAD), consisting of well-defined low-level daily activities intended as a benchmark for algorithm comparison particularly for healthcare scenarios. We briefly review some existing publicly available datasets and compare them with USC-HAD. We describe the wearable sensors used and details of dataset construction. We use high-precision well-calibrated sensing hardware such that the collected data is accurate, reliable, and easy to interpret. The goal is to make the dataset and research based on it repeatable and extendible by others. Mi Zhang 0002, Alexander A. Sawchuk |
UbiComp | 1 |
| 2012 | Sparse representation for motion primitive-based human activity modeling and recognition using wearable sensors
Mi Zhang 0002, Wenyao Xu, Alexander A. Sawchuk, Majid Sarrafzadeh |
ICPR | 1 |