Yuzhe Yang 0003

dblp:213/0962-3 · DBLP profile ↗
← Back
19ranked-venue papers
10as first author
11since 2021 · last 2025
0000-0002-7634-8295ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 6 first-author · 9 since 2021Computer networks · 5 · 4 first-authorGraphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Scaling Wearable Foundation Models
abstract
Wearable sensors have become ubiquitous thanks to a variety of health tracking features. The resulting continuous and longitudinal measurements from everyday life generate large volumes of data. However, making sense of these observations for scientific and actionable insights is non-trivial. Inspired by the empirical success of generative modeling, where large neural networks learn powerful representations from vast amounts of text, image, video, or audio data, we investigate the scaling properties of wearable sensor foundation models across compute, data, and model size. Using a dataset of up to 40 million hours of in-situ heart rate, heart rate variability, accelerometer, electrodermal activity, skin temperature, and altimeter per-minute data from over 165,000 people, we create LSM, a multimodal foundation model built on the largest wearable-signals dataset with the most extensive range of sensor modalities to date. Our results establish the scaling laws of LSM for tasks such as imputation, interpolation and extrapolation across both time and sensor modalities. Moreover, we highlight how LSM enables sample-efficient downstream learning for tasks including exercise and activity recognition.
Girish Narayanswamy, Xin Liu 0034, Kumar Ayush, Yuzhe Yang 0003, Xuhai Xu, Shun Liao, Jake Garrison, Shyam A. Tailor, Jacob E. Sunshine, Yun Liu 0013, Tim Althoff, Shri Narayanan, Pushmeet Kohli, Jiening Zhan, Mark Malhotra, Shwetak N. Patel, Samy Abdel-Ghaffar, Daniel McDuff
ICLR4
2025 RADAR: Benchmarking Language Models on Imperfect Tabular Data
abstract
Language models (LMs) are increasingly being deployed to perform autonomous data analyses. However, their data awareness—the ability to recognize, reason over, and appropriately handle data artifacts such as missing values, outliers, and logical inconsistencies—remains underexplored. These artifacts are especially common in real-world tabular data and, if mishandled, can significantly compromise the validity of analytical conclusions. To address this gap, we present RADAR, a benchmark for systematically evaluating data-aware reasoning on tabular data. We develop a framework to simulate data artifacts via programmatic perturbations to enable targeted evaluation of model behavior. RADAR comprises 2,980 table-query pairs, grounded in real-world data spanning 9 domains and 5 data artifact types. In addition to evaluating artifact handling, RADAR systematically varies table size to study how reasoning performance holds when increasing table size. Our evaluation reveals that, despite decent performance on tables without data artifacts, frontier models degrade significantly when data artifacts are introduced, exposing critical gaps in their capacity for robust, data-aware analysis. Designed to be flexible and extensible, RADAR supports diverse perturbation types and controllable table sizes, offering a valuable resource for advancing tabular reasoning.
Ken Gu, Zhihan Zhang 0002, Kate Lin, Yuwei Zhang 0001, Akshay Paruchuri, Hong Yu 0001, Mehran Kazemi, Kumar Ayush, A. Ali Heydari, Maxwell A. Xu, Yun Liu 0013, Ming-Zher Poh, Yuzhe Yang 0003, Mark Malhotra, Shwetak N. Patel, Hamid Palangi, Xuhai Xu, Daniel McDuff, Tim Althoff, Xin Liu 0034
NeurIPS13
2025 SensorLM: Learning the Language of Wearable Sensors
abstract
We present SensorLM, a family of sensor-language foundation models that enable wearable sensor data understanding with natural language. Despite its pervasive nature, aligning and interpreting sensor data with language remains challenging due to the lack of paired, richly annotated sensor-text descriptions in uncurated, real-world wearable data. We introduce a hierarchical caption generation pipeline designed to capture statistical, structural, and semantic information from sensor data. This approach enabled the curation of the largest sensor-language dataset to date, comprising over 59.7 million hours of data from more than 103,000 people. Furthermore, SensorLM extends prominent multimodal pretraining architectures (e.g., CLIP, CoCa) and recovers them as specific variants within a generic architecture. Extensive experiments on real-world tasks in human activity analysis and healthcare verify the superior performance of SensorLM over state-of-the-art in zero-shot recognition, few-shot learning, and cross-modal retrieval. SensorLM also demonstrates intriguing capabilities including scaling behaviors, label efficiency, sensor captioning, and zero-shot generalization to unseen tasks. Code is available at https://github.com/Google-Health/consumer-health-research/tree/main/sensorlm.
Yuwei Zhang 0001, Kumar Ayush, Siyuan Qiao, A. Ali Heydari, Girish Narayanswamy, Maxwell A. Xu, Ahmed Metwally 0002, Jinhua Xu, Jake Garrison, Xuhai Xu, Tim Althoff, Yun Liu 0013, Pushmeet Kohli, Jiening Zhan, Mark Malhotra, Shwetak N. Patel, Cecilia Mascolo, Xin Liu 0034, Daniel McDuff, Yuzhe Yang 0003
NeurIPS20
2024 BigSmall: Efficient Multi-Task Learning for Disparate Spatial and Temporal Physiological Measurements
abstract
Understanding of human visual perception has historically inspired the design of computer vision architectures. As an example, perception occurs at different scales both spatially and temporally, suggesting that the extraction of salient visual information may be made more effective by attending to specific features at varying scales. Visual changes in the body, due to physiological processes, also occur at varying scales and with modality-specific characteristic properties. Inspired by this, we present BigSmall, an efficient architecture for physiological and behavioral measurement. We present the first joint camera-based facial action, cardiac, and pulmonary measurement model. We propose a multi-branch network with wrapping temporal shift modules that yields efficiency gains and accuracy on par with task-optimized methods. We observe that fusing low-level features leads to suboptimal performance, but that fusing high level features enables efficiency gains with negligible losses in accuracy. We experimentally validate that BigSmall significantly reduces computational cost while achieving comparable results on multiple physiological measurement tasks simultaneously with a unified model.
Girish Narayanswamy, Yuzhe Yang 0003, Chengqian Ma, Xin Liu 0034, Daniel McDuff, Shwetak N. Patel
WACV3
2024 Towards long-tailed, multi-label disease classification from chest X-ray: Overview of the CXR-LT challenge
Gregory Holste, Yiliang Zhou, Song Wang 0026, Ajay Jaiswal, Mingquan Lin, Sherry Zhuge, Yuzhe Yang 0003, Dongkyun Kim, Trong-Hieu Nguyen Mau, Minh-Triet Tran, Jaehyup Jeong, Wongi Park, Jong Bin Ryu, Feng Hong 0004, Arsh Verma, Yosuke Yamagishi, Hyeryeong Seo, Myungjoo Kang, Leo A. Celi, Zhiyong Lu, Ronald M. Summers, George Shih, Zhangyang Wang, Yifan Peng 0002
Medical Image Anal.7
2023 SimPer: Simple Self-Supervised Learning of Periodic Targets
Yuzhe Yang 0003, Xin Liu 0034, Silviu Borac, Dina Katabi, Ming-Zher Poh, Daniel McDuff
ICLR1
2023 Change is Hard: A Closer Look at Subpopulation Shift
abstract
Machine learning models often perform poorly on subgroups that are underrepresented in the training data. Yet, little is understood on the variation in mechanisms that cause subpopulation shifts, and how algorithms generalize across such diverse shifts at scale. In this work, we provide a fine-grained analysis of subpopulation shift. We first propose a unified framework that dissects and explains common shifts in subgroups. We then establish a comprehensive benchmark of 20 state-of-the-art algorithms evaluated on 12 real-world datasets in vision, language, and healthcare domains. With results obtained from training over 10,000 models, we reveal intriguing observations for future progress in this space. First, existing algorithms only improve subgroup robustness over certain types of shifts but not others. Moreover, while current algorithms rely on group-annotated validation data for model selection, we find that a simple selection criterion based on worst-class accuracy is surprisingly effective even without any group information. Finally, unlike existing works that solely aim to improve worst-group accuracy (WGA), we demonstrate the fundamental tradeoff between WGA and other important metrics, highlighting the need to carefully choose testing metrics. Code and data are available at: https://github.com/YyzHarry/SubpopBench.
Yuzhe Yang 0003, Haoran Zhang 0003, Dina Katabi, Marzyeh Ghassemi
ICML1
2023 Rank-N-Contrast: Learning Continuous Representations for Regression
abstract
Deep regression models typically learn in an end-to-end fashion without explicitly emphasizing a regression-aware representation. Consequently, the learned representations exhibit fragmentation and fail to capture the continuous nature of sample orders, inducing suboptimal results across a wide range of regression tasks. To fill the gap, we propose Rank-N-Contrast (RNC), a framework that learns continuous representations for regression by contrasting samples against each other based on their rankings in the target space. We demonstrate, theoretically and empirically, that RNC guarantees the desired order of learned representations in accordance with the target orders, enjoying not only better performance but also significantly improved robustness, efficiency, and generalization. Extensive experiments using five real-world regression datasets that span computer vision, human-computer interaction, and healthcare verify that RNC achieves state-of-the-art performance, highlighting its intriguing properties including better data efficiency, robustness to spurious targets and data corruptions, and generalization to distribution shifts.
Kaiwen Zha, Jeany Son, Yuzhe Yang 0003, Dina Katabi
NeurIPS4
2022 Targeted Supervised Contrastive Learning for Long-Tailed Recognition
abstract
Real-world data often exhibits long tail distributions with heavy class imbalance, where the majority classes can dominate the training process and alter the decision bound-aries of the minority classes. Recently, researchers have in-vestigated the potential of supervised contrastive learning for long-tailed recognition, and demonstrated that it provides a strong performance gain. In this paper, we show that while supervised contrastive learning can help improve performance, past baselines suffer from poor uniformity brought in by imbalanced data distribution. This poor uni-formity manifests in samples from the minority class having poor separability in the feature space. To address this problem, we propose targeted supervised contrastive learning (TSC), which improves the uniformity of the feature distribution on the hypersphere. TSC first generates a set of targets uniformly distributed on a hypersphere. It then makes the features of different classes converge to these distinct and uniformly distributed targets during training. This forces all classes, including minority classes, to main-tain a uniform distribution in the feature space, improves class boundaries, and provides better generalization even in the presence of long-tail data. Experiments on multi-ple datasets show that TSC achieves state-of-the-art performance on long-tailed recognition tasks.
Tianhong Li, Yuan Yuan 0002, Lijie Fan, Yuzhe Yang 0003, Rogério Feris, Piotr Indyk, Dina Katabi
CVPR5
2022 On Multi-Domain Long-Tailed Recognition, Imbalanced Domain Generalization and Beyond
Yuzhe Yang 0003, Hao Wang 0014, Dina Katabi
ECCV (20)1
2021 Delving into Deep Imbalanced Regression
abstract
Real-world data often exhibit imbalanced distributions, where certain target values have significantly fewer observations. Existing techniques for dealing with imbalanced data focus on targets with categorical indices, i.e., different classes. However, many tasks involve continuous targets, where hard boundaries between classes do not exist. We define Deep Imbalanced Regression (DIR) as learning from such imbalanced data with continuous targets, dealing with potential missing data for certain target values, and generalizing to the entire target range. Motivated by the intrinsic difference between categorical and continuous label space, we propose distribution smoothing for both labels and features, which explicitly acknowledges the effects of nearby targets, and calibrates both label and learned feature distributions. We curate and benchmark large-scale DIR datasets from common real-world tasks in computer vision, natural language processing, and healthcare domains. Extensive experiments verify the superior performance of our strategies. Our work fills the gap in benchmarks and techniques for practical imbalanced regression problems. Code and data are available at: https://github.com/YyzHarry/imbalanced-regression.
Yuzhe Yang 0003, Kaiwen Zha, Ying-Cong Chen, Hao Wang 0014, Dina Katabi
ICML1
2020 When NAS Meets Robustness: In Search of Robust Architectures Against Adversarial Attacks
abstract
Recent advances in adversarial attacks uncover the intrinsic vulnerability of modern deep neural networks. Since then, extensive efforts have been devoted to enhancing the robustness of deep networks via specialized learning algorithms and loss functions. In this work, we take an architectural perspective and investigate the patterns of network architectures that are resilient to adversarial attacks. To obtain the large number of networks needed for this study, we adopt one-shot neural architecture search, training a large network for once and then finetuning the sub-networks sampled therefrom. The sampled architectures together with the accuracies they achieve provide a rich basis for our study. Our ''robust architecture Odyssey'' reveals several valuable observations: 1) densely connected patterns result in improved robustness; 2) under computational budget, adding convolution operations to direct connection edge is effective; 3) flow of solution procedure (FSP) matrix is a good indicator of network robustness. Based on these observations, we discover a family of robust architectures (RobNets). On various datasets, including CIFAR, SVHN, Tiny-ImageNet, and ImageNet, RobNets exhibit superior robustness performance to other widely used architectures. Notably, RobNets substantially improve the robust accuracy (~5% absolute gains) under both white-box and black-box attacks, even with fewer parameter numbers. Code is available at https://github.com/gmh14/RobNets.
Yuzhe Yang 0003, Rui Xu 0014, Ziwei Liu 0002, Dahua Lin
CVPR2
2020 Harnessing Structures for Value-Based Planning and Reinforcement Learning
Yuzhe Yang 0003, Guo Zhang 0006, Zhi Xu 0001, Dina Katabi
ICLR1
2019 ME-Net: Towards Effective Adversarial Robustness with Matrix Estimation
abstract
Deep neural networks are vulnerable to adversarial attacks. The literature is rich with algorithms that can easily craft successful adversarial examples. In contrast, the performance of defense techniques still lags behind. This paper proposes ME-Net, a defense method that leverages matrix estimation (ME). In ME-Net, images are preprocessed using two steps: first pixels are randomly dropped from the image; then, the image is reconstructed using ME. We show that this process destroys the adversarial structure of the noise, while re-enforcing the global structure in the original image. Since humans typically rely on such global structures in classifying images, the process makes the network mode compatible with human perception. We conduct comprehensive experiments on prevailing benchmarks such as MNIST, CIFAR-10, SVHN, and Tiny-ImageNet. Comparing ME-Net with state-of-the-art defense mechanisms shows that ME-Net consistently outperforms prior techniques, improving robustness against both black-box and white-box attacks.
Yuzhe Yang 0003, Guo Zhang 0006, Zhi Xu 0001, Dina Katabi
ICML1
2019 ImgSensingNet: UAV Vision Guided Aerial-Ground Air Quality Sensing System
abstract
Given the increasingly serious air pollution problem, air quality index (AQI) monitoring in urban areas has drawn considerable attention. This paper presents ImgSensingNet, a vision guided aerial-ground sensing system, for air quality monitoring and forecasting by the fusion of haze images taken by the unmanned-aerial-vehicle (UAV) and the AQI data collected by an on-ground wireless sensor network. Specifically, ImgSensingNet first leverages the computer vision technique to tell the AQI scale in different regions from the haze images, where haze-relevant features and a deep convolutional neural network (CNN) are designed for direct learning between haze images and corresponding AQI scale. Based on the learnt AQI scale, ImgSensingNet determines whether to wake up on-ground wireless sensors for small-scale AQI monitoring and inference, which can greatly reduce the energy consumption of the system. An entropy-based model is employed for accurate real-time AQI estimation at un-measured locations and future air quality distribution forecasting. We implement and evaluate ImgSensingNet on two university campuses since Feb. 2018, and has collected 17,630 photos and 2.6 millions of AQI data samples. Experimental results confirm that ImgSensingNet can achieve high estimation accuracy while greatly reduce the battery consumption, compared to other state-of-the-art AQI monitoring approaches.
Yuzhe Yang 0003, Zhiwen Hu, Kaigui Bian, Lingyang Song
INFOCOM1
2018 Optimal Trajectory Planning of Drones for 3D Mobile Sensing
abstract
Mobile sensing is challenging in 3D space, as there are many inaccessible places where people rarely venture. Unmanned aerial vehicle (UAV), commonly known as drone, has greatly extended the scope of mobile sensing in 3D space, and pushed forward a variety of 3D mobile sensing applications, such as aerial photo- or video-graphy, 3D wireless signal survey, and air quality monitoring. However, the short battery life of drones has largely restricted the wide adoption of these applications. In this paper, we study the trajectory planning problem for optimizing the flight route in a given sensing space. We first divide the 3D space into an infinite three-dimensional network of observation locations (OLs), and model the sensing scope as a finite subgraph of 3D OL network. We formulate the problem as finding the optimal trajectory in the sensing scope. We propose an algorithm that finds trajectory in each divided 3D grid of the sensing scope by generating a nearly optimal dominating path, and finding the minimum dominating set in the dominating path. Then, we concatenate obtained trajectories in 3D grids to a nearly optimal trajectory in the sensing scope. Experimental results show that the proposed algorithm takes 24% less time to complete sensing the given space, and during the battery life it can cover 19% more sensing scope, than existing solutions.
Yuzhe Yang 0003, Yuanxing Zhang, Kaigui Bian, Lingyang Song, Pengpeng Qiao, Zhetao Li
GLOBECOM2
2018 Sensor Deployment Recommendation for 3D Fine-Grained Air Quality Monitoring Using Semi-Supervised Learning
abstract
Driven by the increasingly serious air pollution problem, the monitoring of fine-grained air quality index (AQI) in urban areas has drawn considerable attention. In this paper, we design a novel algorithm to recommend the placement of sensors for energy- efficient AQI monitoring in urban three-dimensional (3D) space. Specifically, we first propose an entropy- based semi-supervised learning (ESSL) model to estimate the AQI distribution of unobserved locations, using the sparse historical spatial-temporal data and other features, including 3D coordinates, wind speed and weather conditions. Based on ESSL, we then design an entropy minimization ranking (EMR) algorithm to recommend the best sensor locations for AQI monitoring. Through the emulation on a fine-grained AQI dataset, the results demonstrate our scheme can provide energy- efficient solutions by using the least number of sensors to achieve higher accuracy than other existing approaches.
Yuzhe Yang 0003, Kaigui Bian, Lingyang Song, Zhu Han 0001
ICC1
2018 Real-Time Profiling of Fine-Grained Air Quality Index Distribution Using UAV Sensing
abstract
Given significant air pollution problems, air quality index (AQI) monitoring has recently received increasing attention. In this paper, we design a mobile AQI monitoring system boarded on the unmanned-aerial-vehicles, called ARMS, to efficiently build fine-grained AQI maps in real-time. Specifically, we first propose the Gaussian plume model on the basis of the neural network (GPM-NN), to physically characterize the particle dispersion in the air. Based on GPM-NN, we propose a battery efficient and adaptive monitoring algorithm to monitor AQI at the selected locations and construct an accurate AQI map with the sensed data. The proposed adaptive monitoring algorithm is evaluated in two typical scenarios, a 2-D open space like a roadside park, and a 3-D space like a courtyard inside a building. The experimental results demonstrate that our system can provide higher prediction accuracy of AQI with GPM-NN than other existing models, while greatly reducing the power consumption with the adaptive monitoring algorithm.
Yuzhe Yang 0003, Kaigui Bian, Lingyang Song, Zhu Han 0001
IEEE Internet Things J.1
2017 Arms: A Fine-Grained 3D AQI Realtime Monitoring System by UAV
abstract
Recently, mobile devices have been used to carry sensors to monitor air quality index (AQI), and help construct an AQI map in 2-dimensional (2D) areas. In this paper, we design a novel 3-dimensional (3D) AQI monitoring system, called Arms (AQI realtime monitoring system), to efficiently build realtime fine-grained 3D AQI maps, with the help of unmanned-aerial- vehicles (UAVs). Based on the data monitored by Arms, a novel dispersion model, namely Adaptive Gaussian Plume Model (AGPM) is proposed to predict the distribution of AQI. Moreover, the adaptive monitoring techniques, i.e., complete and optimized monitoring, are designed to effectively produce and maintain realtime AQI maps, while greatly reducing the measurement efforts. Experimental results verify that Arms can provide higher predicting accuracy of AQI with the proposed AGPM than other existing models. In addition, the whole system's battery consumption can be greatly reduced.
Yuzhe Yang 0003, Kaigui Bian, Lingyang Song, Zhu Han 0001
GLOBECOM1