Baichen Yang

dblp:271/9958 · DBLP profile ↗
← Back
8ranked-venue papers
1as first author
6since 2021 · last 2026
0000-0002-3031-2668ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 6 · 1 first-author · 5 since 2021Systems, architecture and hardware · 2 · 1 since 2021
YearPublicationVenuePosition
2026 MeatSpec-G: Generalized Low-Cost Spectral Imaging for Ubiquitous Meat Fraud Inspection
abstract
Meat adulteration is a significant problem that can pose health risks economic losses to consumers. Current detection methods are hindered by high costs, limited capabilities, or time-consuming sample preparation, making them only accessible in laboratory tests and can not protect the safety of end-users. This paper introduces MeatSpec, a low-cost and user-friendly system for detecting meat adulteration using spectral imaging, to move the adulteration inspection out of laboratories. MeatSpec employs a multispectral camera to reduce costs while quickly capturing spectral images, but this leads to a decrease in spectral resolution and coverage. To solve this challenge, the system uses spectral reconstruction technology and innovative designs tailored for meat adulteration detection. This includes involving adulteration-related prior information during the reconstruction training phase and incorporating contrastive learning to enlarge the distances among reconstructed samples belonging to various adulteration types. Additionally, we devise distinct feature extractors for different bands based on characteristics of the reconstructed spectra and employ knowledge distillation to mitigate error in full-band reconstructed spectra while capturing features related to adulteration. Further, we extend our system to MeatSpec-G to improve its generalizability to varied adulteration conditions and unknown adulterants. To achieve this, we first propose a feature alignment-based training scheme to reduce the feature gap among samples of diverse concentrations and admixture patterns. Then, we propose a cascaded open-set recognition framework that decouples uncertainty quantification and anomaly feature discrimination, to address the limitations of softmax confidence in detecting distribution shifts and reconstruction artifacts. Experimental evaluations on 347 paired spectral images demonstrate that our system achieves a 91.06% accuracy in detecting multiple adulteration types, merely 7.78% inferior to the expensive professional solution, yet 21.58% superior to the baseline at the same price point. Moreover, our system can generalize to achieve an 88.89% detection accuracy in unknown adulteration conditions with a 27.78% improvement, and an 83.33% detection accuracy for unknown adulterants.
Yinan Zhu, Haiyan Hu 0003, Baichen Yang, Hua Kang, Shanwen Chen, Qianyi Huang, Qian Zhang 0001
IEEE Trans. Mob. Comput.3
2025 EasySpiro: Assessing Lung Function via Arbitrary Exhalations on Commodity Earphones
abstract
Conventional pulmonary function tests (PFTs) are important but costly. Hence, prior research has proposed IoT sensor-based solutions to facilitate cost-efficient, at-home PFT. However, these solutions require the subject to perform maximal exhalations, a task often challenging without supervision, compromising test accuracy. In response to this challenge, this study introduces EasySpiro that, for the first time, uses non-maximal exhalations to measure PFT indicators. This is challenging since PFT indicators are only defined for maximal exhalations, and there are no guidelines to derive them from submaximal exhalations. To address that, we observe that pulmonary deficiencies affect all types of breathing, where the underlying pulmonary deficiency should be the same under different breathing efforts. Leveraging this insight, we design a reconstruction model to predict the ideal maximal breathing patterns based on submaximal ones and utilize these reconstructions for PFT. Furthermore, since the body dynamics reflect the exhalation effort, we use self-supervised learning techniques to encode body dynamics into breathing effort representations to guide the reconstruction process. We integrate these designs into earphones with microphones to measure breathing patterns and IMUs to measure body dynamics. We collaborate with a hospital and develop a dataset from 50 patients with various diseases to evaluate EasySpiro's performance, which shows an accurate prediction of PFT indicators based on non-maximal exhalations with an error rate of 7%. In addition, we open-source the collected dataset to encourage future research.
Wentao Xie 0001, Baichen Yang, Yanbin Gong, Jin Zhang 0001, Shifang Yang, Qian Zhang 0001
MobiCom3
2025 FreshSpec: Sashimi Freshness Monitoring With Low-Cost Multispectral Devices
abstract
Monitoring sashimi freshness,i.e., histamine levels, in showcases poses a critical challenge for sushi restaurants and fresh food stores. Current histamine monitoring methods involve labor-intensive chemical experiments or expensive devices, making affordable on-site monitoring difficult. This paper proposes FreshSpec, a low-cost and automatic spectral imaging system capable of precisely monitoring histamine levels in sashimi with minimal human intervention. The low concentration of histamine, combined with the potential for other ingredients to mask its spectral characteristics, complicates precise histamine level predictions using coarse or redundant spectral data from low-cost devices. To address this issue, FreshSpec employs an innovative feature- wise spectral reconstruction (SR) framework that effectively eliminates irrelevant and redundant data while preserving critical histamine-related spectral features. Specifically, we redefine the SR reconstruction target by utilizing features derived from the encoder of the spectral foundation model that is enhanced to focus on histamine-related spectral features. Furthermore, inspired by the monotonic accumulation properties of histamine over time, we propose a histamine regression model with unsupervised continual adaptation to new sashimi samples during practical deployment. Experimental results from 240 samples of salmon, tuna, and snapper demonstrate that FreshSpec achieves an R2 of 0.9319 and an RMSE of 3.101 mg/100 g, comparable to laboratory spectral imaging systems, while outperforming baseline schemes with a 46.95% RMSE reduction and a 0.1631 R2 improvement.
Yinan Zhu, Haiyan Hu 0003, Baichen Yang, Qianyi Huang, Qian Zhang 0001
IEEE Trans. Mob. Comput.3
2024 MeatSpec: Enabling Ubiquitous Meat Fraud Inspection through Consumer-Level Spectral Imaging
abstract
Meat adulteration is a significant problem that can pose health risks economic losses to consumers. Current detection methods are hindered by high costs, limited capabilities, or time-consuming sample preparation, making them only accessible in laboratory tests and can not protect the safety of end-users. This paper introduces MeatSpec, a low-cost and user-friendly system for detecting meat adulteration using spectral imaging, to move the adulteration inspection out of laboratories. MeatSpec employs a multispectral camera to reduce costs while quickly capturing spectral images, but this leads to a decrease in spectral resolution and coverage. To solve this challenge, the system uses spectral reconstruction technology and innovative designs tailored for meat adulteration detection. This includes involving adulteration-related prior information during the reconstruction training phase and incorporating contrastive learning to enlarge the distances among reconstructed samples belonging to various adulteration types. Additionally, we devise distinct feature extractors for different bands based on characteristics of the reconstructed spectra and employ knowledge distillation to mitigate error in full-band reconstructed spectra while capturing features related to adulteration. Experimental evaluations on 347 paired spectral images demonstrate that our system achieves a 91.06% accuracy in detecting multiple adulteration types, merely 7.78% inferior to the expensive professional solution, yet 21.58% superior to the baseline at the same price point.
Haiyan Hu 0003, Yinan Zhu, Baichen Yang, Hua Kang, Shanwen Chen, Qian Zhang 0001
MobiCom3
2023 PDAssess: A Privacy-preserving Free-speech based Parkinson's Disease Daily Assessment System
abstract
In-time disease assessment is essential to better customize the medication scheme and improve the quality of life for chronic diseases like Parkinson's disease (PD). Toward the inconvenience problem in current clinical assessment practice, mobile sensing solutions based on detecting Parkinson's vocal changes are proposed. However, current solutions either can only achieve binary disease detection task or require patients to perform specific speaking tasks, which is not effective and practical for disease stage assessment in daily scenario. Moreover, most of existing solutions do not take speech privacy into consideration. In this work, we present PDAssess, a free speech-based daily assessment system that can perform 4-stage Parkinson's disease assessment in a privacy-preserving manner. We observe that current solutions did not fully leverage the rich information embedded in free speech due to the linguistic content variations, and therefore leverage a pre-trained automatic speech recognition (ASR) model to achieve a content variation-aware feature-extraction. In order to distinguish subtle stage-wise differences, we design a novel attention-based neural network architecture with a customized loss function for disease assessment task. Towards the potential privacy leakage problem, we design a Split Learning-based framework with pseudo-labeling and local domain adversarial training to better preserve speech content privacy. We collaborate with a medical center and evaluate the performance of PDAssess on real-world speech data collected from 50 PD subjects and 50 healthy subjects. The evaluation result shows that PDAssess can perform 4-stage PD assessment with an average person-wise F1 score of 89.1% and voice sample-wise F1 score of 75.1%.
Baichen Yang, Qingyong Hu, Wentao Xie 0001, Qian Zhang 0001
SenSys1
2021 Citadel: Protecting Data Privacy and Model Confidentiality for Collaborative Learning
abstract
Many organizations own data but have limited machine learning expertise (data owners). On the other hand, organizations that have expertise need data from diverse sources to train truly generalizable models (model owners). With the advancement of machine learning (ML) and its growing awareness, the data owners would like to pool their data and collaborate with model owners, such that both entities can benefit from the obtained models. In such a collaboration, the data owners want to protect the privacy of its training data, while the model owners desire the confidentiality of the model and the training method that may contain intellectual properties. Existing private ML solutions, such as federated learning and split learning, cannot simultaneously meet the privacy requirements of both data and model owners.
Chengliang Zhang, Junzhe Xia, Baichen Yang, Huancheng Puyang, Wei Wang 0030, Ruichuan Chen, Istemi Ekin Akkus, Paarijaat Aditya, Feng Yan 0001
SoCC3
2020 DIESEL: A Dataset-Based Distributed Storage and Caching System for Large-Scale Deep Learning Training
abstract
We observe three problems in existing storage and caching systems for deep-learning training (DLT) tasks: (1) accessing a dataset containing a large number of small files takes a long time, (2) global in-memory caching systems are vulnerable to node failures and slow to recover, and (3) repeatedly reading a dataset of files in shuffled orders is inefficient when the dataset is too large to be cached in memory. Therefore, we propose DIESEL, a dataset-based distributed storage and caching system for DLT tasks. Our approach is via a storage-caching system co-design. Firstly, since accessing small files is a metadata-intensive operation, DIESEL decouples the metadata processing from metadata storage, and introduces metadata snapshot mechanisms for each dataset. This approach speeds up metadata access significantly. Secondly, DIESEL deploys a task-grained distributed cache across the worker nodes of a DLT task. This way node failures are contained within each DLT task. Furthermore, the files are grouped into large chunks in storage, so the recovery time of the caching system is reduced greatly. Thirdly, DIESEL provides chunk-based shuffle so that the performance of random file access is improved without sacrificing training accuracy. Our experiments show that DIESEL achieves a linear speedup on metadata access, and outperforms an existing distributed caching system in both file caching and file reading. In real DLT tasks, DIESEL halves the data access time of an existing storage system, and reduces the training time by hours without changing any training code.
Lipeng Wang 0004, Songgao Ye, Baichen Yang, Youyou Lu, Hequan Zhang, Shengen Yan, Qiong Luo 0001
ICPP3
2020 RepBun: Load-Balanced, Shuffle-Free Cluster Caching for Structured Data
abstract
Cluster caching systems increasingly store structured data objects in the columnar format. However, these systems routinely face the imbalanced load that significantly impairs the I/O performance. Existing load-balancing solutions, while effective for reading unstructured data objects, fall short in handling columnar data. Unlike unstructured data that can only be read through a full-object scan, columnar data supports direct query of specific columns with two distinct access patterns: (1) columns have the heavily skewed popularity, and (2) hot columns are likely accessed together in a query job. Based on these two access patterns, we propose an effective load-balancing solution for structured data. Our solution, which we call RepBun, groups hot columns into a bundle. It then copies multiple replicas of the column bundle and stores them uniformly across servers. We show that RepBun achieves improved load balancing with reduced memory overhead, while avoiding data shuffling between cache servers. We implemented RepBun atop Alluxio, a popular in-memory distributed storage, and evaluate its performance through EC2 deployment against the TPC-H benchmark work-load. Experimental results show that RepBun outperforms the existing load-balancing solutions with significantly shorter read latency and faster query completion.
Minchen Yu, Yinghao Yu, Yunchuan Zheng, Baichen Yang, Wei Wang 0030
INFOCOM4