Kaiyuan Hou

dblp:118/8687 · DBLP profile ↗
← Back
15ranked-venue papers
5as first author
14since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 9 · 4 first-author · 9 since 2021Systems, architecture and hardware · 4 · 3 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 EmbodiedFly: Embodied LLM Agent with an Autonomous Reconfigurable Drone
abstract
Large Language Models (LLMs) have shown immense human-like capabilities for reasoning and generating digital content. However, their ability to freely sense, interact, and actuate the physical domain remains significantly limited due to three fundamental challenges: (1) physical environments require specialized sensors for different tasks, yet deploying dedicated sensors for each application is impractical; (2) events and objects of interest are often localized to small areas within large spaces, making them difficult to detect with static sensor networks; and (3) foundation models need flexible actuation capabilities to meaningfully interact with the physical world. To bridge this gap, we introduce EmbodiedFly, an embodied LLM agent combining a foundation model pipeline with a reconfigurable drone platform to observe, understand, and interact with the physical world. Our co-design approach features (1) a FM orchestration framework connecting multiple LLMs, VLMs, and an open-set object detection model; (2) a novel image segmentation technique that identifies task-relevant areas; and (3) a custom drone platform that autonomously reconfigures with appropriate sensors and actuators based on commands from the FM orchestration framework. Through real-world deployments, we demonstrate that EmbodiedFly completes diverse physical tasks with up to \(85\%\) higher success rates compared to traditional approaches leveraging static deployments.
Kaiyuan Hou, Junxi Xia, Stephen Xia, Xiaofan Jiang 0001
ACM Trans. Internet Things2
2025 FlexiFly: Interfacing the Physical World with Foundation Models Empowered by Reconfigurable Drone Systems
abstract
Foundation models (FM) have shown immense human-like capabilities for generating digital media. However, foundation models that can freely sense, interact, and actuate the physical domain is far from being realized. This is due to 1) requiring dense deployments of sensors to fully cover and analyze large spaces, while 2) events often being localized to small areas, making it difficult for FMs to pinpoint relevant areas of interest relevant to the current task. We propose FlexiFly, a platform that enables FMs to "zoom in" and analyze relevant areas with higher granularity to better understand the physical environment and carry out tasks. FlexiFly accomplishes by introducing 1) a novel image segmentation technique that aids in identifying relevant locations and 2) a modular and reconfigurable sensing and actuation drone platform that FMs can actuate to "zoom in" with relevant sensors and actuators. We demonstrate through real smart home deployments that FlexiFly enables FMs and LLMs to complete diverse tasks up to 85% more successfully. FlexiFly is critical step towards FMs and LLMs that can naturally interface with the physical world.
Junxi Xia, Kaiyuan Hou, Stephen Xia, Xiaofan Jiang 0001
SenSys3
2024 Improving On-Device LLMs' Sensory Understanding with Embedding Interpolations
abstract
Large Language Models (LLMs) have shown significant potential in performing inferences on various tasks using heterogeneous sensors with minimal human intervention. Despite their promise, challenges such as high inference overhead and limitations on resource-constrained edge devices remain. Additionally, model hallucinations, particularly those arising from cognitive biases when interpreting numerical data, hinder performance. This work introduces a novel technique, embedding interpolation, to enhance LLMs' understanding of sensor measurements and mitigate inference overhead on edge devices. By computing embeddings through pre-computed boundary embeddings instead of directly from the input, we improve efficiency and accuracy. The effective-ness of this approach is demonstrated through visualizations with image generation models.
Kaiyuan Hou, Yunqi Guo, Heming Fu, Hongkai Chen 0001, Zhenyu Yan 0002, Guoliang Xing, Xiaofan Jiang 0001
MobiCom1
2024 Connecting Foundation Models with the Physical World using Reconfigurable Drone Agents
abstract
Foundation models excel in tasks such as content generation, zero-shot classifications, and reasoning. However, they struggle with sensing, interacting, and actuating in the physical world due to their dependence on limited sensors and actuators in providing timely contextual information or physical interactions. This reliance restricts the system's adaptability and coverage. To address these issues and create an embodied AI with foundation models (FMs), we introduce Embodied Reconfigurable Drone Agent (EmbodiedRDA). EmbodiedRDA features a custom drone platform that can autonomously swap payloads to reconfigure itself with a diverse list of sensors and actuators. We designed FM agents to instruct the drone to equip itself with appropriate physical modules, analyze sensor data, make decisions, and control the drone's actions. This enables the system to perform a variety of tasks in dynamic physical environments, bridging the gap between the digital and physical worlds.
Kaiyuan Hou, Junxi Xia, Stephen Xia, Xiaofan Jiang 0001
MobiCom2
2024 h5bench: A unified benchmark suite for evaluating HDF5 I/O performance on pre-exascale platforms
abstract
Summary Parallel I/O is a critical technique for moving data between compute and storage subsystems of supercomputers. With massive amounts of data produced or consumed by compute nodes, high‐performant parallel I/O is essential. I/O benchmarks play an important role in this process; however, there is a scarcity of I/O benchmarks representative of current workloads on HPC systems. Toward creating representative I/O kernels from real‐world applications, we have created h5bench , a set of I/O kernels that exercise hierarchical data format version 5 (HDF5) I/O on parallel file systems in numerous dimensions. Our focus on HDF5 is due to the parallel I/O library's heavy usage in various scientific applications running on supercomputing systems. The various tests benchmarked in the h5bench suite include I/O operations (read and write), data locality (arrays of basic data types and arrays of structures), array dimensionality (one‐dimensional arrays, two‐dimensional meshes, three‐dimensional cubes), I/O modes (synchronous and asynchronous). In this paper, we present the observed performance of h5bench executed along several of these dimensions on existing supercomputers (Cori and Summit) and pre‐exascale platforms (Perlmutter, Theta, and Polaris). h5bench measurements can be used to identify performance bottlenecks and their root causes and evaluate I/O optimizations. As the I/O patterns of h5bench are diverse and capture the I/O behaviors of various HPC applications, this study will be helpful to the broader supercomputing and I/O community.
Jean Luca Bez, Houjun Tang, M. Scot Breitenfeld, Huihuo Zheng, Wei-keng Liao, Kaiyuan Hou, Zanhua Huang, Surendra Byna
Concurr. Comput. Pract. Exp.6
2023 ARSteth: Enabling Home Self-Screening with AR-Assisted Intelligent Stethoscopes
abstract
The stethoscope is one of the most important diagnostic tools used by healthcare professionals, through a process called auscultation, to screen patients for abnormalities of the heart and lungs. While there are digital stethoscopes on the market which ease this process, it still takes years of training to properly use these devices to listen for abnormal sounds within the body. We present ARSteth, an intelligent stethoscope platform that improves the accessibility of stethoscopes for the general population, allowing anyone to perform auscultation in the comfort of their own homes. Our platform utilizes a combination of augmented reality (AR), acoustic intelligence, and human-machine interaction to dynamically guide users on where to place the stethoscope on different parts of the body (auscultation points), through visual and audio cues. Through user studies, we show that ARSteth, on average, can guide users within 13.2 mm from optimal auscultation points marked by licensed physicians in 13.09 seconds for each auscultation point. By guiding users towards more effective auscultation points, make preventative health screening more accessible and effective for everyone we are able to achieve higher confidence on classifying heart murmurs.
Kaiyuan Hou, Stephen Xia, Emily Bejerano, Junyi Wu 0004, Xiaofan Jiang 0001
IPSN1
2023 Anemoi: A Low-cost Sensorless Indoor Drone System for Automatic Mapping of 3D Airflow Fields
abstract
Mapping 3D airflow fields is important for many HVAC, industrial, medical, and home applications. However, current approaches are expensive and time-consuming. We present Anemoi, a sub-$100 drone-based system for autonomously mapping 3D airflow fields in indoor environments. Anemoi leverages the effects of airflow on motor control signals to estimate the magnitude and direction of wind at any given point in space. We introduce an exploration algorithm for selecting optimal waypoints that minimize overall airflow estimation uncertainty. We demonstrate through microbenchmarks and real deployments that Anemoi is able to estimate wind speed and direction with errors up to 0.41 m/s and 25.1° lower than the existing state of the art and map 3D airflow fields with an average RMS error of 0.73 m/s.
Stephen Xia, Charuvahan Adhivarahan, Kaiyuan Hou, Jingping Nie, Eugene Wu 0002, Karthik Dantu, Xiaofan Jiang 0001
MobiCom4
2023 I/O in WRF: A Case Study in Modern Parallel I/O Techniques
abstract
Large-scale parallel applications can face significant I/O performance bottlenecks, making efficient I/O crucial. This work presents a comparative study of several parallel I/O implementations in the Weather Research and Forecasting model, including PnetCDF blocking and non-blocking I/O options, netCDF4, HDF5 Log VOL, and ADIOS. For I/O methods creating files in a canonical data layout, PnetCDF's non-blocking option offers up to 2x improvement over its blocking option and up to 4.5x over HDF5 via netCDF4, demonstrating the effectiveness of the write request aggregation technique. The HDF5 Log VOL outperforms ADIOS with a 4x improvement in write performance when creating files in the log layout, although both require non-negligible time to convert the file back to canonical order for post-run analysis. From these results we extract some observations that can guide I/O strategies for modern parallel codes.
Zanhua Huang, Kaiyuan Hou, Ankit Agrawal 0001, Alok N. Choudhary, Robert B. Ross, Wei-keng Liao
SC2
2022 A Low-Cost In-situ System for Continuous Multi-Person Fever Screening
abstract
With the recent societal impact of COVID-19, companies and government agencies alike have turned to thermal camera based skin temperature sensing technology to help screen for fever. However, the cost and deployment restrictions limit the wide use of these thermal sensing technologies. In this work, we present SIFTER, a low-cost system based on a RGB-thermal camera for continuous fever screening of multiple people. This system detects and tracks heads in the RGB and thermal domains and constructs thermal heat map models for each tracked person, and classifies people as having or not having fever. SIFTER can obtain key temperature features of heads in-situ at a distance and produce fever screening predictions in real-time, significantly improving screening through-put while minimizing disruption to normal activities. In our clinic deployment, SIFTER measurement error is within 0.4°F at 2 meters and around 0.6°F at 3.5 meters. In comparison, most infrared thermal scanners on the market costing several thousand dollars have around 1°F measurement error measured within 0.5 meters. SIFTER can achieve 100% true positive rate with 22.5% false positive rate without requiring any human interaction, greatly outperforming our baseline [1], which sees a false positive rate of 78.5%.
Kaiyuan Hou, Peter Wei, Chenye Yang, Hengjiu Kang, Stephen Xia, Teresa Spada, Andrew Rundle, Xiaofan Jiang 0001
IPSN1
2022 A modular and reconfigurable sensing and actuation platform for smarter environments and drones: demo abstract
abstract
There has been an immense growth in sensors, actuators, and smart devices in recent years, which enable us to better sense, actuate, and understand the physical world. Despite this growth, we have yet to achieve fully intelligent environments. This is, in part, due to the large number of different organizations creating smart devices with proprietary technologies and communication protocols that are not compatible with each other and require significant engineering to incorporate and adapt to specific applications. In this work, we present an easy-to-install and low-cost embedded platform that allows users to rapidly configure a mixture of sensors and actuators. The system is based on the commonly-used Raspberry Pi ecosystem, easily configurable, and does not require users to have prior knowledge of programming, which allows anyone, regardless of background, to use. We also introduce a battery-powered wireless extension module that is suitable for mobile drone applications, where a chord-powered Raspberry Pi is not suitable. We demonstrate the impact our system has on enabling drones with flexible sensing modalities and creating smarter environments by integrating our platform into a variety of intelligent home applications.
Avik Dhupar, Kaiyuan Hou, Stephen Xia, Xiaofan Jiang 0001
MobiSys4
2022 AI Stethoscope for Home Self-Diagnosis with AR Guidance
abstract
Cardiopulmonary ailments are a major cause of mortality. Stethoscopes are one of the most important tools that healthcare professionals use to screen patients for a variety of ailments, especially those related to the heart and lungs. Despite the growth of digital stethoscopes on the market, it takes years of training to properly use stethoscopes to listen for abnormal sounds within the body. In this demonstration, we present an intelligent stethoscope platform that makes stethoscopes more accessible to the general population. Our platform utilizes augmented reality (AR) to provide real-time guidance on where to properly place the stethoscope on the body, enabling the general population to screen themselves for ailments.
Kaiyuan Hou, Stephen Xia, Junyi Wu 0004, Emily Bejerano, Xiaofan Jiang 0001
SenSys1
2022 A case study on parallel HDF5 dataset concatenation for high energy physics data analysis
Sunwoo Lee 0001, Kaiyuan Hou, Kewei Wang 0002, Saba Sehrish, Marc F. Paterno, Jim Kowalkowski, Quincey Koziol, Robert B. Ross, Ankit Agrawal 0001, Alok N. Choudhary, Wei-keng Liao
Parallel Comput.2
2021 Supporting Data Compression in PnetCDF
abstract
Recently, the dramatic increase of the data amounts drives up the demand for data compression among HPC applications. Although many file systems and I/O middlewares have incorporated compression features, few high-level parallel I/O libraries support data compression due to the challenges of achieving scalable performance on HPC systems. This paper presents the design and implementation of the variable compression feature in the Parallel NetCDF library. Our design employs the same concept of chunking used by the HDF5 library, but we focus on enabling I/O aggregation across multiple requests to address the challenges on performance and scalability. We evaluate our solution using the I/O kernel of real-world scientific applications and analyze the impacts of data compression on parallel I/O performance. Our result suggests that handling multiple requests at once can significantly improve the parallel I/O performance on chunked and compressed data.
Kaiyuan Hou, Qiao Kang, Sunwoo Lee 0001, Ankit Agrawal 0001, Alok N. Choudhary, Wei-keng Liao
IEEE BigData1
2021 Optimizing Performance of Parallel I/O Accesses to Non-contiguous Blocks in Multiple Array Variables
abstract
Accessing non-contiguous blocks in multiple array variables is a challenging I/O pattern for parallel applications to obtain good I/O performance. High-level I/O libraries such as HDF5 allow users to implement this pattern conveniently, but users have observed significant performance bottlenecks in the two-phase I/O implementation of MPI-IO. Recent studies have advanced the two-phase I/O performance by novel communication algorithms, but such improvements still have limitations. Two-phase I/O has to faithfully process inputs from high-level I/O libraries, so that implementation overheads can accumulate for improper usage of high-level I/O libraries. In this paper, we propose approaches for efficient usage of high-level I/O libraries that can circumvent major collective I/O overheads. We adopt a multi-dataset implementation of HDF5 dataset I/O to aggregate non-contiguous requests for array blocks and provide corresponding parameter assignment strategies. These approaches reduce the overheads caused by communication straggler effects in two-phase I/O. We show that our proposed methods can improve the parallel I/O performance up to 8× on two supercomputing systems for the HDF5 implementations of an I/O kernel extracted from climate simulation code compared with its baseline implementations.
Qiao Kang, M. Scot Breitenfeld, Kaiyuan Hou, Wei-keng Liao, Robert B. Ross, Surendra Byna
IEEE BigData3
2020 Improving MPI Collective I/O for High Volume Non-Contiguous Requests With Intra-Node Aggregation
abstract
Two-phase I/O is a well-known strategy for implementing collective MPI-IO functions. It redistributes I/O requests among the calling processes into a form that minimizes the file access costs. As modern parallel computers continue to grow into the exascale era, the communication cost of such request redistribution can quickly overwhelm collective I/O performance. This effect has been observed from parallel jobs that run on multiple compute nodes with a high count of MPI processes on each node. To reduce the communication cost, we present a new design for collective I/O by adding an extra communication layer that performs request aggregation among processes within the same compute nodes. This approach can significantly reduce inter-node communication contention when redistributing the I/O requests. We evaluate the performance and compare it with the original two-phase I/O on Cray XC40 parallel computers (Theta and Cori) with Intel KNL and Haswell processors. Using I/O patterns from two large-scale production applications and an I/O benchmark, we show our proposed method effectively reduces the communication cost and hence maintains the scalability for a large number of processes.
Qiao Kang, Sunwoo Lee 0001, Kaiyuan Hou, Robert B. Ross, Ankit Agrawal 0001, Alok N. Choudhary, Wei-keng Liao
IEEE Trans. Parallel Distributed Syst.3