Yu-Lin Wei

dblp:162/5525 · DBLP profile ↗
← Back
14ranked-venue papers
4as first author
5since 2021 · last 2025
0000-0002-6518-028XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 8 · 3 first-authorGraphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer networks
9 papers
Wireless sensing and localization · 52% Physical-layer communications · 32% Internet of things and sensor networks · 10%
Artificial intelligence
3 papers
Probabilistic and Bayesian machine learning · 72% Kernel, tree and ensemble methods · 26% Image recognition and object detection · 2%
Human-computer interaction and pervasive computing
6 papers
Wearable and physiological sensing · 20% Interaction techniques and input · 20% Immersive interaction · 20%
Theoretical computer science
2 papers
Mathematical optimization · 100%
Computer graphics and multimedia
2 papers
Rendering · 72% Computational photography and imaging · 28%

Topics — the 19 heaviest of 25, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes › gaussian process
gaussian process regression
1.622025
Kernel Learning for Sample Constrained Black-Box Optimization · AAAI 2025
Sample-Constrained Black Box Optimization for Audio Personalization · AAAI 2024
Mathematical optimization
black-box optimization
1.622025
Kernel Learning for Sample Constrained Black-Box Optimization · AAAI 2025
Sample-Constrained Black Box Optimization for Audio Personalization · AAAI 2024
Machine learning › Kernel, tree and ensemble methods › kernel methods
kernel learning
0.912025
Kernel Learning for Sample Constrained Black-Box Optimization · AAAI 2025
Rendering
neural radiance fields
0.912025
Can NeRFs "See" without Cameras? · NeurIPS 2025
Physical-layer communications › optical wireless communication
visible light communication
0.832017
Demo: CELLI: Indoor Positioning using Polarized Sweeping Light Beams · MobiSys 2017
CELLI: Indoor Positioning Using Polarized Sweeping Light Beams · MobiSys 2017
RollingLight: Enabling Line-of-Sight Light-to-Camera Communications · MobiSys 2015
Wireless sensing and localization
indoor localization
0.842020
Demo: CELLI: Indoor Positioning using Polarized Sweeping Light Beams · MobiSys 2017
CELLI: Indoor Positioning Using Polarized Sweeping Light Beams · MobiSys 2017
Ear-AR: indoor acoustic augmented reality on earphones · MobiCom 2020
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes › gaussian process › sparse gaussian process
sparse gaussian process regression
0.812024
Sample-Constrained Black Box Optimization for Audio Personalization · AAAI 2024
Wireless sensing and localization › indoor localization
visible light positioning
0.622017
Demo: CELLI: Indoor Positioning using Polarized Sweeping Light Beams · MobiSys 2017
CELLI: Indoor Positioning Using Polarized Sweeping Light Beams · MobiSys 2017
Immersive interaction › augmented reality
audio augmented reality
0.412020
Ear-AR: indoor acoustic augmented reality on earphones · MobiCom 2020
Wearable and physiological sensing › earable sensing
earphone-based sensing
0.412020
EarSense: earphones as a teeth activity sensor · MobiCom 2020
Interaction techniques and input
gesture input
0.412020
EarSense: earphones as a teeth activity sensor · MobiCom 2020
Wireless sensing and localization › indoor localization
acoustic localization
0.412020
Voice localization using nearby wall reflections · MobiCom 2020
Physical-layer communications › signal processing for communications › array signal processing
direction-of-arrival estimation
0.412020
Voice localization using nearby wall reflections · MobiCom 2020
Wireless sensing and localization › tracking
indoor tracking
0.312018
Augmenting Indoor Inertial Tracking with Polarized Light · MobiSys 2018
Internet of things and sensor networks › motion sensing › inertial sensing
orientation estimation
0.312017
LiCompass: Extracting orientation from polarized light · INFOCOM 2017
Accessibility and assistive technology › assistive technology
hearing aid personalization
0.312025
Kernel Learning for Sample Constrained Black-Box Optimization · AAAI 2025
Wireless sensing and localization
wifi sensing
0.212016
Location-Independent WiFi Action Recognition via Vision-based Methods · ACM Multimedia 2016
Physical-layer communications › optical wireless communication
light-to-camera communication
0.212015
RollingLight: Enabling Line-of-Sight Light-to-Camera Communications · MobiSys 2015
Human-AI interaction
voice assistants
0.112020
Voice localization using nearby wall reflections · MobiCom 2020

Methods — techniques the papers use, named apart from their topics

variational autoencoder · 2.6kernel learning · 2.6gaussian process · 2.6surrogate modeling · 2.3gaussian process regression · 2.3neural radiance field · 1.7implicit neural representation · 1.7IMU fusion · 1.1reverse triangulation · 0.9microphone array · 0.93d audio · 0.9polarized sweeping light beams · 0.6polarization-modulated signals · 0.6LCD projection · 0.6vibration sensing · 0.4acoustic signal analysis · 0.4acoustic calibration · 0.4polarized light · 0.3
YearPublicationVenuePosition
2025 Kernel Learning for Sample Constrained Black-Box Optimization
abstract
Black box optimization (BBO) focuses on optimizing unknown functions in high-dimensional spaces. In many applications, sampling the unknown function is expensive, imposing a tight sample budget.Ongoing work is making progress on reducing the sample budget by learning the shape/structure of the function, known as kernel learning. We propose a new method to learn the kernel of a Gaussian Process. Our idea is to create a continuous kernel space in the latent space of a variational autoencoder, and run an auxiliary optimization to identify the best kernel. Results show that the proposed method, Kernel Optimized Blackbox Optimization (KOBO), outperforms state of the art by estimating the optimal at considerably lower sample budgets. Results hold not only across synthetic benchmark functions but also in real applications. We show that a hearing aid may be personalized with fewer audio queries to the user, or a generative model could converge to desirable images from limited user ratings.
Rajalaxmi Rajagopalan, Yu-Lin Wei, Romit Roy Choudhury
AAAI2
2025 Estimating Multi-chirp Parameters using Curvature-guided Langevin Monte Carlo
abstract
This paper considers the problem of estimating chirp parameters from a noisy mixture of chirps. While a rich body of work exists in this area, challenges remain when extending these techniques to chirps of higher order polynomials. We formulate this as a non-convex optimization problem and propose a modified Langevin Monte Carlo (LMC) sampler that exploits the average curvature of the objective function to reliably find the minimizer. Results show that our Curvature-guided LMC (CG-LMC) algorithm is robust and succeeds even in low SNR regimes, making it viable for practical applications.
Sattwik Basu, Debottam Dutta, Yu-Lin Wei, Romit Roy Choudhury
ICASSP3
2025 Can NeRFs "See" without Cameras?
abstract
Neural Radiance Fields (NeRFs) have been remarkably successful at synthesizing novel views of 3D scenes by optimizing a volumetric scene function. This scene function models how optical rays bring color information from a 3D object to the camera pixels. Radio frequency (RF) or audio signals can also be viewed as a vehicle for delivering information about the environment to a sensor. However, unlike camera pixels, an RF/audio sensor receives a mixture of signals that contain many environmental reflections (also called “multipath”). Is it still possible to infer the environment using such multipath signals? We show that with redesign, NeRFs can be taught to learn from multipath signals, and thereby “see” the environment. As a grounding application, we aim to infer the indoor floorplan of a home from sparse WiFi measurements made at multiple locations inside the home. Although a difficult inverse problem, our implicitly learnt floorplans look promising, and enables forward applications, such as indoor signal prediction and basic ray tracing.
Chaitanya Amballa, Yu-Lin Wei, Sattwik Basu, Zhijian Yang, Mehmet Ergezer, Romit Roy Choudhury
NeurIPS2
2024 Sample-Constrained Black Box Optimization for Audio Personalization
abstract
We consider the problem of personalizing audio to maximize user experience. Briefly, we aim to find a filter h*, which applied to any music or speech, will maximize the user’s satisfaction. This is a black-box optimization problem since the user’s satisfaction function is unknown. Substantive work has been done on this topic where the key idea is to play audio samples to the user, each shaped by a different filter hi, and query the user for their satisfaction scores f(hi). A family of “surrogate” functions is then designed to fit these scores and the optimization method gradually refines these functions to arrive at the filter ˆh* that maximizes satisfaction. In certain applications, we observe that a second type of querying is possible where users can tell us the individual elements h*[j] of the optimal filter h*. Consider an analogy from cooking where the goal is to cook a recipe that maximizes user satisfaction. A user can be asked to score various cooked recipes (e.g., tofu fried rice) or to score individual ingredients (say, salt, sugar, rice, chicken, etc.). Given a budget of B queries, where a query can be of either type, our goal is to find the recipe that will maximize this user’s satisfaction. Our proposal builds on Sparse Gaussian Process Regression (GPR) and shows how a hybrid approach can outperform any one type of querying. Our results are validated through simulations and real world experiments, where volunteers gave feedback on music/speech audio and were able to achieve high satisfaction levels. We believe this idea of hybrid querying opens new problems in black-box optimization and solutions can benefit other applications beyond audio personalization.
Rajalaxmi Rajagopalan, Yu-Lin Wei, Romit Roy Choudhury
AAAI2
2021 Angle-of-Arrival (AoA) Factorization in Multipath Channels
abstract
This paper considers the problem of estimating K angle of arrivals (AoA) using an array of M > K microphones. We assume the source signal is human voice, hence unknown to the receiver. Moreover, the signal components that arrive over K spatial paths are strongly correlated since they are delayed copies of the same source signal. Past works have successfully extracted the AoA of the direct path, or have assumed specific types of signals/channels to derive the subsequent (multipath) AoAs. Our method builds on the core observation that signals from multiple AoAs embed predictable delay-structures in them, which can be factorized through iterative alignment and cancellation. Simulation results show median AoA errors of < 4° for the first 3 AoAs. Real-world measurements, from a circular microphone array similar to Amazon Echo, show modest degradations. We believe the ability to infer even K = 3 AoAs can be helpful to various sensing and localization applications.
Yu-Lin Wei, Romit Roy Choudhury
ICASSP1
2020 EarSense: earphones as a teeth activity sensor
abstract
This paper finds that actions of the teeth, namely tapping and sliding, produce vibrations in the jaw and skull. These vibrations are strong enough to propagate to the edge of the face and produce vibratory signals at an earphone. By re-tasking the earphone speaker as an input transducer - a software modification in the sound card - we are able to sense teeth-related gestures across various models of ear/headphones. In fact, by analyzing the signals at the two earphones, we show the feasibility of also localizing teeth gestures, resulting in a human-to-machine interface. Challenges range from coping with weak signals, distortions due to different teeth compositions, lack of timing resolution, spectral dispersion, etc. We address these problems with a sequence of sensing techniques, resulting in the ability to detect 6 distinct gestures in real-time. Results from 18 volunteers exhibit robustness, even though our system - EarSense - does not depend on per-user training. Importantly, EarSense also remains robust in the presence of concurrent user activities, like walking, nodding, cooking and cycling. Our ongoing work is focused on detecting teeth gestures even while music is being played in the earphone; once that problem is solved, we believe EarSense could be even more compelling.
Jay Prakash, Zhijian Yang, Yu-Lin Wei, Haitham Hassanieh, Romit Roy Choudhury
MobiCom3
2020 Voice localization using nearby wall reflections
abstract
Voice assistants such as Amazon Echo (Alexa) and Google Home use microphone arrays to estimate the angle of arrival (AoA) of the human voice. This paper focuses on adding user localization as a new capability to voice assistants. For any voice command, we desire Alexa to be able to localize the user inside the home. The core challenge is two-fold: (1) accurately estimating the AoAs of multipath echoes without the knowledge of the source signal, and (2) tracing back these AoAs to reverse triangulate the user's location.
Sheng Shen 0002, Daguan Chen, Yu-Lin Wei, Zhijian Yang, Romit Roy Choudhury
MobiCom3
2020 Ear-AR: indoor acoustic augmented reality on earphones
abstract
This paper aims to use modern earphones as a platform for acoustic augmented reality (AAR). We intend to play 3D audio-annotations in the user's ears as she moves and looks at AAR objects in the environment. While companies like Bose and Microsoft are beginning to release such capabilities, they are intended for outdoor environments. Our system aims to explore the challenges indoors, without requiring any infrastructure deployment. Our core idea is two-fold. (1) We jointly use the inertial sensors (IMUs) in earphones and smartphones to estimate a user's indoor location and gazing orientation. (2) We play 3D sounds in the earphones and exploit the human's responses to (re)calibrate errors in location and orientation. We believe this fusion of IMU and acoustics is novel, and could be an important step towards indoor AAR. Our system, Ear-AR, is tested on 7 volunteers invited to an AAR exhibition - like a museum - that we set up in our building's lobby and lab. Across 60 different test sessions, the volunteers browsed different subsets of 24 annotated objects as they walked around. Results show that Ear-AR plays the correct audio-annotations with good accuracy. The user-feedback is encouraging and points to further areas of research and applications.
Zhijian Yang, Yu-Lin Wei, Sheng Shen 0002, Romit Roy Choudhury
MobiCom2
2018 Augmenting Indoor Inertial Tracking with Polarized Light
abstract
Inertial measurement unit (IMU) has long suffered from the problem of integration drift, where sensor noises accumulate quickly and cause fast-growing tracking errors. Existing methods for calibrating IMU tracking either require human in the loop, or need energy-consuming cameras, or suffer from coarse tracking granularity. We propose to augment indoor inertial tracking by reusing existing indoor luminaries to project a static light polarization pattern in the space. This pattern is imperceptible to human eyes and yet through a polarizer, it becomes detectable by a color sensor, and thus can serve as fine-grained optical landmarks that constrain and correct IMU's integration drift and boost tracking accuracy. Exploiting the birefringence optical property of transparent tapes -- a low-cost and easily-accessible material -- we realize the polarization pattern by simply adding to existing light cover a thin polarizer film with transparent tape stripes glued atop. When fusing with IMU sensor signals, the light pattern enables robust, accurate and low-power motion tracking. Meanwhile, our approach entails low deployment overhead by reusing existing lighting infrastructure without needing an active modulation unit. We build a prototype of our light cover and the sensing unit using off-the-shelf components. Experiments show 4.3 cm median error for 2D tracking and 10 cm for 3D tracking, as well as its robustness in diverse settings.
Tian Zhao 0003, Yu-Lin Wei, Wei-Nin Chang, Changxi Zheng, Hsin-Mu Tsai, Kate Ching-Ju Lin
MobiSys2
2017 LiCompass: Extracting orientation from polarized light
abstract
Accurate orientation information is the key in many applications, ranging from map reconstruction with crowdsourcing data, location data analytics, to accurate indoor localization. Many existing solutions rely on noisy magnetic and inertial sensor data, leading to limited accuracy, while others leverage multiple, dense anchor points to improve the accuracy, requiring significant deployment efforts. This paper presents LiCompass, the first system that enables a commodity camera to accurately estimate the object orientation using just a single optical anchor. Our key idea is to allow a camera to observe varying intensity level of polarized light when it is in different orientations and, hence, perform estimation directly from image pixel intensity. As the estimation relies only on pixel intensity, instead of the location of the anchor in an image, the system performs reliably at long distance, with low resolution images, and with large perspective distortion. LiCompass' core designs include an elaborate optical anchor design and a series of signal processing techniques based on trigonometric properties, which extend the range of orientation estimation to full 360 degrees. Our prototype evaluation shows that LiCompass produces very accurate estimates with median errors of merely 2.5 degrees at 5 meters and 7.4 degrees at 2.5 meters with an irradiance angle of 55 degrees.
Yu-Lin Wei, Hsin-I Wu, Han-Chung Wang, Hsin-Mu Tsai, Kate Ching-Ju Lin, Rayana Boubezari, Hoa Le Minh, Zabih Ghassemlooy
INFOCOM1
2017 CELLI: Indoor Positioning Using Polarized Sweeping Light Beams
abstract
Existing visible light positioning (VLP) systems leverage the high image resolution of a receiving camera to support high positioning accuracy. However, power-hungry cameras might not always be applicable for many scenarios, such as smart factories, in which small objects require to be accurately localized and tracked. In this paper, we introduce CELLI, an indoor VLP system that only uses a single luminary as the transmitter and requires only a simple light sensor to achieve an extremely high accuracy with centimeter-level error. The key idea is to provide the spatial resolution capability from the transmitter instead of the receiver, so that the complexity of the receiver can be minimized. In particular, a small LCD is installed at the transmitter to project a large number of narrow and interference-free polarized light beams to different spatial cells. A receiving light sensor identifies its located cell by detecting the unique polarization-modulated signals projected to that cell. CELLI further incorporates a number of novel designs to overcome the technical challenges such as reducing the positioning latency, which is typically limited by the long optical response time of an LCD, and transforming a cell coordinate to the global 3D position using only a single light. We have prototyped our design using off-the-shelf optical and electrical components, and experimentally shown that CELLI achieves a median 3D positioning error less than 11.8 cm and a median 2D positioning error to less than 2.7 cm.
Yu-Lin Wei, Chang-Jung Huang, Hsin-Mu Tsai, Kate Ching-Ju Lin
MobiSys1
2017 Demo: CELLI: Indoor Positioning using Polarized Sweeping Light Beams
abstract
Indoor positioning enables location-based services for a wide range of commercial applications [4]. Existing visible light positioning (VLP) systems [5, 7] leverage the high image resolution of a receiving camera to support high positioning accuracy. However, power-hungry cameras are not desirable in many scenarios, e.g., smart factories, where small objects need to be accurately located and tracked. In this demo, we introduce CELLI, an indoor VLP system that only uses a single luminary as the transmitter and requires only a simple light sensor to achieve high accuracy with centimeter-level error. The key idea is to provide the spatial resolution capability from the transmitter instead of the receiver, so that the complexity of the receiver can be minimized. In particular, a small Liquid Crystal Display (LCD) is installed at the transmitter to project a large number of narrow and interference-free polarized light beams to different spatial cells. A receiving light sensor identifies its located cell by detecting the unique polarization-modulated signals projected to that cell, as shown in Fig. 1. CELLI further incorporates several novel designs to overcome the technical challenges such as reducing the positioning latency, which is typically limited by the long optical response time of an LCD, and transforming a cell coordinate to the global 3D position using only a single light. We have prototyped our design using off-the-shelf optical and electronic components, and experimentally shown that CELLI achieves a median 3D positioning error less than 12 cm and a median 2D error less than 2.7 cm.
Yu-Lin Wei, Chang-Jung Huang, Hsin-Mu Tsai, Kate Ching-Ju Lin
MobiSys1
2016 Location-Independent WiFi Action Recognition via Vision-based Methods
abstract
Due to the characteristics of ubiquity, non-occlusion, privacy preservation of WiFi, many researchers have devoted to human action recognition using WiFi signals. As demonstrated in [1], Channel State Information (CSI), a fine-grained information capturing the properties of WiFi signal propagation, could be transformed into images for achieving a promising accuracy on action recognition via vision-based methods. However, from the experimental results shown in [1], the CSI is usually location dependent, which affects the recognition performance if signals are recorded in different places.
Jen-Yin Chang, Kuan-Ying Lee, Yu-Lin Wei, Kate Ching-Ju Lin, Winston H. Hsu
ACM Multimedia3
2015 RollingLight: Enabling Line-of-Sight Light-to-Camera Communications
abstract
Recent literatures have demonstrated the feasibility and applicability of light-to-camera communications. They either use this new technology to realize specific applications, e.g., localization, by sending repetitive signal patterns, or consider non-line-of-sight scenarios. We however notice that line-of-sight light-to-camera communications has a great potential because it provides a natural way to enable visual association, i.e., visually associating the received information with the transmitter's identity. Such capability benefits broader applications, such as augmented reality, advertising, and driver assistance systems. Hence, this paper designs, implements, and evaluates RollingLight, a line-of-sight light-to-camera communication system that enables a light to talk to diverse off-the-shelf rolling shutter cameras. To boost the data rate and enhance reliability, RollingLight addresses the following practical challenges. First, its demodulation algorithm allows cameras with heterogeneous sampling rates to accurately decode high-order frequency modulation in real-time. Second, it incorporates a number of designs to resolve the issues caused by inherently unsynchronized light-to-camera channels. We have built a prototype of RollingLight with USRP-N200, and also implemented a real system with Arduino Mega 2560, both tested with a range of different camera receivers. We also implement a real iOS application to examine our real-time decoding capability. The experimental results show that, even to serve commodity cameras with a large variety of frame rates, RollingLight can still deliver a throughput of 11.32 bytes per second.
Hui-Yu Lee, Hao-Min Lin, Yu-Lin Wei, Hsin-I Wu, Hsin-Mu Tsai, Kate Ching-Ju Lin
MobiSys3