EDBT 2026 Demo / reviewers in the wild / expert
Kai Huang 0007
dblp:86/489-7
· DBLP profile ↗
11ranked-venue papers
5as first author
10since 2021 · last 2025
0000-0003-2569-2309ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 9 · 4 first-author · 8 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Modality Plug-and-Play: Runtime Modality Adaptation in LLM-Driven Autonomous Mobile SystemsabstractMultimodal reasoning by LLMs is critical to autonomous mobile systems, but the growing diversity of input data modalities prevents incorporating all modalities into LLMs. Instead, only the useful modalities should be adaptively involved at runtime, based on the current environmental contexts and task requirements. Existing work on runtime modality adaptation uses fixed connections between data encoders and LLM's input layer, but results in high training costs and ineffective cross-modal interaction. In this paper, we present MPnP, a new modality adaptation technique that connects data encoders to a flexible set of last LLM blocks and makes such latent connections fully trainable at runtime. Evaluation results show that MPnP has high compute and data efficiency, with 3.7× FLOPs reduction and 30% memory usage reduction compared to best baselines. It requires only few hundreds of training samples at runtime, and completes modality adaptation within few minutes on weak devices. Kai Huang 0007, Xiangyu Yin 0002, Heng Huang 0001, Wei Gao 0006 |
MobiCom | 1 |
| 2024 | Towards Green AI in Fine-tuning Large Language Models via Adaptive BackpropagationabstractFine-tuning is essential to adapting pre-trained large language models to downstream applications. With the increasing popularity of LLM-enabled applications, fine-tuning has been performed intensively worldwide, incurring a tremendous amount of computing costs that correspond to big carbon footprint and environmental impact. Mitigating such environmental impact directly correlates to reducing the fine-tuning FLOPs. Existing fine-tuning schemes focus on either saving memory or reducing the overhead of computing weight updates, but cannot achieve sufficient FLOPs reduction due to their ignorance of the training cost in backpropagation. To address this limitation, in this paper we present GreenTrainer, a new technique that minimizes the FLOPs of LLM fine-tuning via adaptive backpropagation, which adaptively selects the most appropriate set of LLM tensors for fine-tuning based on their importance and backpropagation cost in training. Experiment results show that GreenTrainer can save up to 64\% training FLOPs compared to full fine-tuning, without any noticeable accuracy loss. Compared to the existing schemes such as Prefix Tuning and LoRA, GreenTrainer can achieve up to 4\% improvement of model accuracy, with on-par FLOPs reduction. Kai Huang 0007, Hanyun Yin, Heng Huang 0001, Wei Gao 0006 |
ICLR | 1 |
| 2024 | Perceptual-Centric Image Super-Resolution using Heterogeneous Processors on Mobile DevicesabstractImage super-resolution (SR) is widely used on mobile devices to enhance user experience. However, neural networks used for SR are computationally expensive, posing challenges for mobile devices with limited computing power. A viable solution is to use heterogeneous processors on mobile devices, especially the specialized hardware AI accelerators, for SR computations, but the reduced arithmetic precision on AI accelerators can lead to degraded perceptual quality in upscaled images. To address this limitation, in this paper we present SR For Your Eyes (FYE-SR), a novel image SR technique that enhances the perceptual quality of upscaled images when using heterogeneous processors for SR computations. FYE-SR strategically splits the SR model and dispatches different layers to heterogeneous processors, to meet the time constraint of SR computations while minimizing the impact of AI accelerators on image quality. Experiment results show that FYE-SR outperforms the best baselines, improving perceptual image quality by up to 2×, or reducing SR computing latency by up to 5.6× with on-par image quality. Kai Huang 0007, Xiangyu Yin 0002, Tao Gu 0001, Wei Gao 0006 |
MobiCom | 1 |
| 2023 | ElasticTrainer: Speeding Up On-Device Training with Runtime Elastic Tensor SelectionabstractOn-device training is essential for neural networks (NNs) to continuously adapt to new online data, but can be time-consuming due to the device's limited computing power. To speed up on-device training, existing schemes select trainable NN portion offline or conduct unrecoverable selection at runtime, but the evolution of trainable NN portion is constrained and cannot adapt to the current need for training. Instead, runtime adaptation of on-device training should be fully elastic, i.e., every NN substructure can be freely removed from or added to the trainable NN portion at any time in training. In this paper, we present ElasticTrainer, a new technique that enforces such elasticity to achieve the required training speedup with the minimum NN accuracy loss. Experiment results show that ElasticTrainer achieves up to 3.5× more training speedup in wall-clock time and reduces energy consumption by 2×-3× more compared to the existing schemes, without noticeable accuracy loss. Kai Huang 0007, Boyuan Yang 0001, Wei Gao 0006 |
MobiSys | 1 |
| 2023 | PTEase: Objective Airway Examination for Pulmonary Telemedicine using Commodity SmartphonesabstractRemote monitoring and evaluation of pulmonary diseases via tele-medicine are important to disease diagnosis and management, but current telemedicine solutions have limited capability of objectively examining the airway's internal physiological conditions that are crucial to pulmonary disease evaluation. Existing solutions based on smartphone sensing are also limited to externally monitoring breath rates, respiratory events, or lung function. In this paper, we present PTEase, a new system design that addresses these limitations and uses commodity smartphones to examine the airway's internal physiological conditions. PTEase uses active acoustic sensing to measure the internal changes of lower airway caliber, and then leverages machine learning to analyze the sensory data for pulmonary disease evaluation. We implemented PTEase as a smartphone app, and verified its measurement error in lab-controlled settings as <10%. Clinical studies further showed that PTEase reaches 75% accuracy on disease prediction and 11%-15% errors in estimating lung function indices. Given that such accuracy is comparable with that in clinical practice using spirometry, PTEase can be reliably used as an assistive telemedicine tool for disease evaluation and monitoring. Xiangyu Yin 0002, Kai Huang 0007, Erick Forno, Wei Chen 0074, Heng Huang 0001, Wei Gao 0006 |
MobiSys | 2 |
| 2022 | Eavesdropping user credentials via GPU side channels on smartphonesabstractGraphics Processing Unit (GPU) on smartphones is an effective target for hardware attacks. In this paper, we present a new side channel attack on mobile GPUs of Android smartphones, allowing an unprivileged attacker to eavesdrop the user's credentials, such as login usernames and passwords, from their inputs through on-screen keyboard. Our attack targets on Qualcomm Adreno GPUs and investigate the amount of GPU overdraw when rendering the popups of user's key presses of inputs. Such GPU overdraw caused by each key press corresponds to unique variations of selected GPU performance counters, from which these key presses can be accurately inferred. Experiment results from practical use on multiple models of Android smartphones show that our attack can correctly infer more than 80% of user's credential inputs, but incur negligible amounts of computing overhead and network traffic on the victim device. To counter this attack, this paper suggests mitigations of access control on GPU performance counters, or applying obfuscations on the values of GPU performance counters. Boyuan Yang 0001, Ruirong Chen, Kai Huang 0007, Jun Yang 0002, Wei Gao 0006 |
ASPLOS | 3 |
| 2022 | FaceListener: Recognizing Human Facial Expressions via Acoustic Sensing on Commodity HeadphonesabstractFacial expressions are important indicators of user needs that can be used in many interactive computing applications to adapt the system behaviors and settings. Current computing approaches to recognizing human facial expressions, however, either rely on con-tinuous camera recordings that are energy consuming, or require custom sensing hardware that are expensive and difficult to use on commodity systems. In this paper, we present FaceListener, a new sensing system that recognizes human facial expressions by only using commodity headphones. The basic idea of FaceListener is to transform the commodity headphone into an acoustic sensing device, which captures the face skin deformations caused by fa-cial muscle movements with different facial expressions. To ensure the recognition accuracy, FaceListener leverages the knowledge distillation technique to learn the subtle correlation between face skin deformation and the acoustic signal changes. Experiment re-sults over multiple human beings demonstrate that FaceListener can accurately recognize more than 80% of different facial expressions. FaceListener is highly energy efficient, and can well adapt to different headphone models, host systems and user activities. Xingzhe Song, Kai Huang 0007, Wei Gao 0006 |
IPSN | 2 |
| 2022 | Real-time neural network inference on extremely weak devices: agile offloading with explainable AIabstractWith the wide adoption of AI applications, there is a pressing need of enabling real-time neural network (NN) inference on small embedded devices, but deploying NNs and achieving high performance of NN inference on these small devices is challenging due to their extremely weak capabilities. Although NN partitioning and offloading can contribute to such deployment, they are incapable of minimizing the local costs at embedded devices. Instead, we suggest to address this challenge via agile NN offloading, which migrates the required computations in NN offloading from online inference to offline learning. In this paper, we present AgileNN, a new NN offloading technique that achieves real-time NN inference on weak embedded devices by leveraging eXplainable AI techniques, so as to explicitly enforce feature sparsity during the training phase and minimize the online computation and communication costs. Experiment results show that AgileNN's inference latency is >6X lower than the existing schemes, ensuring that sensory data on embedded devices can be timely consumed. It also reduces the local device's resource consumption by >8X, without impairing the inference accuracy. Kai Huang 0007, Wei Gao 0006 |
MobiCom | 1 |
| 2022 | AiFi: AI-Enabled WiFi Interference Cancellation with Commodity PHY-Layer InformationabstractInterference could result in significant performance degradation in WiFi networks. Most existing solutions to interference cancellation require extra RF hardware, which is usually infeasible in many low-power wireless scenarios. In this paper, we present AiFi, a new interference cancellation technique that can be applied to commodity WiFi devices without using any extra RF hardware. The key idea of AiFi is to retrieve knowledge about interference from the locally available physical-layer (PHY) information at the WiFi receiver, including the pilot information (PI) and the channel state information (CSI). AiFi leverages the power of AI to address the possible ambiguity when estimating interference from these PHY information, and incorporates the domain knowledge about WiFi PHY to minimize the neural network complexity. Experiment results show that AiFi can correct 80% of bit errors due to interference and improves the MAC frame reception rate by 18x, with <1ms latency for interference cancellation in each frame. Ruirong Chen, Kai Huang 0007, Wei Gao 0006 |
SenSys | 2 |
| 2022 | Out-Clinic Pulmonary Disease Evaluation via Acoustic Sensing and Multi-Task Learning on Commodity SmartphonesabstractPulmonary diseases, such as asthma and Chronic Obstructive Pulmonary Disease (COPD), constitute a major public health challenge. The disease symptoms, including airway obstruction and inflammation, usually result in changes in airway mechanical properties, such as the caliber and impedance of the airway. To measure such airway properties for disease evaluation and diagnosis purposes, pulmonary function tests (PFT) has been widely adopted. However, most existing PFT systems require expensive and cumbersome hardware that are impossible to be used out of clinic. To allow out-clinic continuous pulmonary disease evaluation, in this paper we present AWARE, a new sensing and AI system that supports accurate and reliable PFT using commodity smartphones. AWARE uses a smartphone to transmit acoustic signals and reconstructs the profile of human airway based on the analysis of reflected acoustic waves captured from the smartphone's microphone. The subject's pulmonary condition is then evaluated by a multi-task learning model that integrates both the airway measurements and the subject's lung function records as the ground truth. Evaluations on 75 human subjects demonstrate that AWARE has the capability to achieve 80% accuracy on distinguishing between humans with healthy pulmonary function and with asthma symptoms. Xiangyu Yin 0002, Kai Huang 0007, Erick Forno, Wei Chen 0074, Heng Huang 0001, Wei Gao 0006 |
SenSys | 2 |
| 2020 | MagHacker: eavesdropping on stylus pen writing via magnetic sensing from commodity mobile devicesabstractStylus pens have been widely used with today's mobile devices to provide a convenient handwriting input method, but also bring a unique security vulnerability that may unveil the user's handwriting contents to a nearby eavesdropper. In this paper, we present MagHacker, a new sensing system that realizes such eavesdropping attack over commodity mobile devices, which monitor and analyze the magnetic field being produced by the stylus pen's internal magnet. MagHacker divides the continuous magnetometer readings into small segments that represent individual letters, and then translates these readings into writing trajectories for letter recognition. Experiment results over realistic handwritings from multiple human beings demonstrate that MagHacker can accurately eavesdrop more than 80% of handwriting with stylus pens, from a distance of 10cm. Only slight degradation in such accuracy is produced when the eavesdropping distance or the handwriting speed increases. MagHacker is highly energy efficient, and can well adapt to different stylus pen models and environmental contexts. Kai Huang 0007, Xingzhe Song, Boyuan Yang 0001, Wei Gao 0006 |
MobiSys | 2 |