Yang Bai 0009

dblp:39/6825-9 · DBLP profile ↗
← Back
8ranked-venue papers
5as first author
5since 2021 · last 2025
0000-0003-4437-9457ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 7 · 5 first-author · 4 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
YearPublicationVenuePosition
2025 SING: Spatial Context in Large Language Model for Next-Gen Wearables
abstract
Integrating spatial context into large language models (LLMs) has the potential to revolutionize human-computer interaction, particularly in wearable devices. In this work, we present a novel system architecture that incorporates spatial speech understanding into LLMs, enabling contextually aware and adaptive applications for wearable technologies. Our approach leverages microstructure-based spatial sensing to extract precise Direction of Arrival (DoA) information using a monaural microphone. To address the lack of existing dataset for microstructure-assisted speech recordings, we synthetically create a dataset by using the LibriSpeech dataset. This spatial information is fused with linguistic embeddings from OpenAI’s Whisper model, allowing each modality to learn complementary contextual representations. The fused embeddings are aligned with the input space of LLaMA-3.2 3B model and fine-tuned with lightweight adaptation technique LoRA to optimize for on-device processing. SING supports spatially-aware automatic speech recognition (ASR), achieving a mean error of 25.72°—a substantial improvement compared to the 88.52° median error in existing work—with a word error rate (WER) of 5.3. SING also supports soundscaping, for example, inference how many people were talking and their directions, with up to 5 people and a median DoA error of 16°. Our system demonstrates superior performance in spatial speech understanding while addressing the challenges of power efficiency, privacy, and hardware constraints, paving the way for advanced applications in augmented reality, accessibility, and immersive experiences.
Ayushi Mishra, Yang Bai 0009, Priyadarshan Narayanasamy, Nakul Garg, Nirupam Roy
ICML2
2022 SPiDR: ultra-low-power acoustic spatial sensing for micro-robot navigation
abstract
This paper presents the design and implementation of SPiDR, an ultra-low-power spatial sensing system for miniature mobile robots. This acoustic sensor produces a cross-sectional map of the field-of-view using only one speaker/microphone pair. While it is challenging to have enough spatial diversity of signal with a single omnidirectional source, we leverage sound's interaction with small structures to create a 3D-printed passive filter, called a stencil, that can project spatially coded signals on a region at a fine granularity. The system receives a linear combination of the reflections from nearby objects and applies a novel power-aware depth-map reconstruction algorithm. The algorithm first estimates the approximate locations of the objects in the scene and then iteratively applies fractional multi-resolution inversion. SPiDR consumes only 10mW of power to generate a depth-map in real-world scenario with over 80% structural similarity score with the scene.
Yang Bai 0009, Nakul Garg, Nirupam Roy
MobiSys1
2022 Ultra-low-power acoustic imaging
abstract
This poster presents the design and implementation of SPiDR, an ultra-low-power acoustic imaging system. This imaging system produces a cross-sectional map of the field-of-view using only one speaker/microphone pair. It leverages the fact that sound's interaction with small structures can project spatially coded signals on a region at a fine granularity. We create a 3D-printed passive filter, called a stencil, that can image the scene with a single omnidirectional source and sensor. With spatially coded signal, the system receives a linear combination of the reflections from nearby objects and applies a novel power-aware depth-map reconstruction algorithm. SPiDR consumes only 10mW of power to generate a depth-map in real-world scenario with over 80% structural similarity score with the scene.
Yang Bai 0009, Nakul Garg, Nirupam Roy
MobiSys1
2021 Owlet: enabling spatial information in ubiquitous acoustic devices
abstract
This paper presents a low-power and miniaturized design for acoustic direction-of-arrival (DoA) estimation and source localization, called Owlet. The required aperture, power consumption, and hardware complexity of the traditional array-based spatial sensing techniques make them unsuitable for small and power-constrained IoT devices. Aiming to overcome these fundamental limitations, Owlet explores acoustic microstructures for extracting spatial information. It uses a carefully designed 3D-printed metamaterial structure that covers the microphone. The structure embeds a direction-specific signature in the recorded sounds. Owlet system learns the directional signatures through a one-time in-lab calibration. The system uses an additional microphone as a reference channel and develops techniques that eliminate environmental variation, making the design robust to noises and multipaths in arbitrary locations of operations. Owlet prototype shows 3.6° median error in DoA estimation and 10cm median error in source localization while using a 1.5cm × 1.3cm acoustic structure for sensing. The prototype consumes less than 100th of the energy required by a traditional microphone array to achieve similar DoA estimation accuracy. Owlet opens up possibilities of low-power sensing through 3D-printed passive structures.
Nakul Garg, Yang Bai 0009, Nirupam Roy
MobiSys2
2021 Microstructure-guided spatial sensing for low-power IoT
abstract
This demonstration presents a working prototype of Owlet, an alternative design for spatial sensing of acoustic signals. To overcome the fundamental limitations in form-factor, power consumption, and hardware requirements with array-based techniques, Owlet explores wave's interaction with acoustic structures for sensing. By combining passive acoustic microstructures with microphones, we envision achieving the same functionalities as microphone and speaker arrays with less power consumption and in a smaller form factor. Our design uses a 3D-printed metamaterial structure over a microphone to introduce a carefully designed spatial signature to the recorded signal. Owlet prototype shows 3.6° median error in Direction-of-Arrival (DoA) estimation and 10 cm median error in source localization while using a 1.5cm × 1.3cm acoustic structure for sensing.
Nakul Garg, Yang Bai 0009, Nirupam Roy
MobiSys2
2020 BatComm: enabling inaudible acoustic communication with high-throughput for mobile devices
abstract
Acoustic communication is an increasingly popular alternative to existing short-range wireless communication technologies for mobile devices, such as NFC and QR codes. Unlike the current standards, there are no requirements for extra hardware, lighting conditions, or Internet connection. However, the audibility and limited throughput of existing studies hinder their deployment on a wide range of applications. In this paper, we aim to redesign acoustic communication mechanism to push the boundary of potential throughput while keeping the inaudibility. Specifically, we propose BatComm, a high-throughput and inaudible acoustic communication system for mobile devices capable of throughput rates 12X higher than contemporary state-of-the-art acoustic communication for mobile devices. We theoretically model the non-linearity of microphone and use orthogonal frequency division multiplexing (OFDM) to transmit data bits over multiple orthogonal channels with an ultrasound frequency carrier. We also design a series of techniques to mitigate interference caused by sources such as the signal's unbalanced frequency response, ambient noise, and unrelated residual signals created through OFDM, amplitude modulation (AM), and related processes. Extensive evaluations under multiple realistic settings demonstrate that our inaudible acoustic communication system can achieve over 47kbps within a 10cm communication range. We also show the possibility of increasing the communication range to room scale (i.e., around 2m) while maintaining high-throughput and inaudibility. Our findings offer a new direction for future inaudible acoustic communication techniques to pursue in emerging mobile and IoT applications.
Yang Bai 0009, Jian Liu 0001, Li Lu 0008, Yingying Chen 0001, Jiadi Yu
SenSys1
2020 Acoustic-based sensing and applications: A survey
Yang Bai 0009, Li Lu 0008, Jerry Q. Cheng, Jian Liu 0001, Yingying Chen 0001, Jiadi Yu
Comput. Networks1
2019 Poster: Inaudible High-throughput Communication Through Acoustic Signals
abstract
In recent decades, countless efforts have been put into the research and development of short-range wireless communication, which offers a convenient way for numerous applications (e.g., mobile payments, mobile advertisement). Regarding the design of acoustic communication, throughput and inaudibility are the most vital aspects, which greatly affect available applications that can be supported and their user experience. Existing studies on acoustic communication either use audible frequency band (e.g., <20kHz) to achieve a relatively high throughput or realize inaudibility using near-ultrasonic frequency band (e.g., 18-20kHz) which however can only achieve limited throughput. Leveraging the non-linearity of microphones, voice commands can be demodulated from the ultrasound signals, and further recognized by the speech recognition systems. In this poster, we design an acoustic communication system, which achieves high-throughput and inaudibility at the same time, and the highest throughput we achieve is over 17x higher than the state-of-the-art acoustic communication systems.
Yang Bai 0009, Jian Liu 0001, Yingying Chen 0001, Li Lu 0008, Jiadi Yu
MobiCom1