VLDB 2026 Research / reviewers in the wild / expert
Yongjian Fu 0004
dblp:09/4381-4
· DBLP profile ↗
21ranked-venue papers
5as first author
21since 2021 · last 2026
0000-0001-8481-2644ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 17 · 5 first-author · 17 since 2021Human-computer interaction and ubiquitous computing · 3 · 3 since 2021Security and privacy · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | EarAuth: Towards Practical Cardiac Vibration Authentication on COTS Wireless Earbuds
Yongjian Fu 0004, Wenpeng Zhu, Yingjun Wu, Hao Pan 0003, Guanbo Wang, Yongheng Deng, Yaoxue Zhang, Ju Ren 0001 |
INFOCOM | 1 |
| 2026 | LiBre: Toward Motion-Resilient Contactless Respiration Monitoring Using Mobile LiDARabstractIn this paper, we present LiBre, a LiDAR-based system for real-time respiration monitoring that remains accurate and continuous even under device motion. LiBre addresses the core challenge of disentangling large-scale device movement from subtle thoracoabdominal motions by integrating three key components: (i) an Object-Centric Feature Extraction module that produces clean, geometrically consistent human point clouds and enables multi-target sensing with minimal environmental interference; (ii) a Contrastive Registration framework that combines standard Iterative Closest Point (ICP) and masked Iterative Closest Point (MICP) to decouple device motion from respiration-induced displacements; and (iii) a Directional Residual Projection strategy that automatically estimates the RoI and projects residual motion along the dominant respiratory axis, eliminating the need for manual annotation. Following a brief autonomous stationary initialization phase to establish the respiratory RoI and optimized weights, we implement a complete prototype and validate its real-time performance. Experiments with 10 participants demonstrate that LiBre achieves respiration monitoring with < 1 BPM error at sensing distances up to 4 m, under device motion speeds up to 30 cm/s, and at orientation angles up to 60°, while supporting multi-person scenarios. The system processes each frame within 120 ms, meeting the requirements for real-time mobile health monitoring in practical applications. Junying Hu, Yongjian Fu 0004, Xinyi Li 0005, Yaoxue Zhang, Ju Ren 0001 |
IEEE Internet Things J. | 2 |
| 2026 | Mobile and Multi-Device Wireless ChargingabstractWireless charging is a cornerstone technology for next-generation mobile and ubiquitous computing. However, its practical deployment has long been constrained by short range, poor flexibility, and lack of support for dynamic multi-device scenarios. In this paper, we propose ChargeX—a system that enables long-range and mobility-resilient wireless charging for multiple small devices. ChargeX pioneers the integration of metasurface-assisted magnetic beamforming, a high-frequency compact transceiver design, and a real-time closed-loop feedback-control mechanism. It further advances the field by introducing a joint optimization framework for dynamically allocating energy across mobile receivers with heterogeneous priorities and spatial-temporal demands. Experimental results demonstrate that it achieves meter-level charging distance, real-time response to device movement, and efficient coordination among multiple receivers, significantly outperforming state-of-the-art prototypes. Bozhong Yu, Yongjian Fu 0004, Ju Ren 0001, Hao Pan 0003, Jeremy Gummeson, Ling Wang 0007, Yaoxue Zhang |
IEEE Trans. Mob. Comput. | 3 |
| 2026 | MagGuard: Detecting Mobile Eavesdropping via Built-In Magnetometers With Contrastive LearningabstractProtecting privacy-sensitive hardware usage on mobile devices is crucial. Although mobile operating systems (OSs) and smartphone manufacturers have set the permission settings, attackers can evade these defenses using covert methods, enabling malicious camera recording, microphone eavesdropping, and screen capture. Electronic devices emit unique yet weak electromagnetic interference (EMI) signals when accessing privacy-sensitive hardware. But, these signals are easily affected by foreground application activities and geomagnetic fluctuations caused by device movement. Our prior work showed that supervised learning can extract EMI features correlated with privacy hardware states from complex magnetometer readings, but it requires substantial labeled data, limiting practical deployment to new device models or OS versions. To eliminate this reliance on labeled data, this paper proposes a multimodal contrastive learning framework that leverages the device's built-in magnetometer and synchronized system logs as dual-modal inputs. Through self-supervised training, the framework can learn the intrinsic associations between EMI features and the operating states of privacy-sensitive hardware. Building on this, we design an EMI-based eavesdropping classifier that can analyze a user device's magnetometer readings offline to detect covert eavesdropping activities. Experimental results show that the proposed method can effectively identify eavesdropping behavior related to access to camera, microphone, and screen recording data. Testing across ten diverse mobile devices achieved an average classification accuracy of 89.1% on Android devices and 88.5% on iOS devices for identifying the specific hardware being eavesdropped upon. Hao Pan 0003, Lanqing Yang, Yongjian Fu 0004, Yi-Chao Chen 0001, Guangtao Xue, Ju Ren 0001 |
IEEE Trans. Mob. Comput. | 3 |
| 2026 | SRDrone: LLM-Driven Self-Refinement for Embodied Drone Task PlanningabstractWe introduceSRDrone, a novel system designed for self-refinement task planning in industrial-grade embodied drones.SRDroneincorporates two key technical contributions: First, it employs a continuous state evaluation methodology to robustly and accurately determine task outcomes and provide explanatory feedback. This approach supersedes conventional reliance on single-frame final-state assessment for continuous, dynamic drone operations. Second,SRDroneimplements a hierarchical Behavior Tree (BT) modification model. This model integrates multi-level BT plan analysis with a constrained strategy space to enable structured reflective learning from experience. Experimental results demonstrate thatSRDroneachieves a 44.87% improvement in Success Rate (SR) over baseline methods. Furthermore, real-world deployment utilizing an experience base optimized through iterative self-refinement attains a 96.25% SR. By embedding adaptive task refinement capabilities within an industrial-grade BT planning framework,SRDroneeffectively integrates the general reasoning intelligence of Large Language Models (LLMs) with the stringent physical execution constraints inherent to embodied drones. Code is available athttps://github.com/ZXiiiC/SRDrone. Tingting Long, Xunhua Dai, Yongjian Fu 0004, Ju Ren 0001, Yaoxue Zhang |
IEEE Trans. Mob. Comput. | 7 |
| 2026 | HyStream: A Hybrid System for Application Streaming via Predictive Delivery and Sequence-Linearized CachingabstractTraditional application delivery requires full local installation, introducing persistent security risks from outdated software and imposing significant download delays. While advances in network bandwidth and latency have made remote content delivery more viable, existing dynamic loading mechanisms, such as network filesystems, often remain constrained by performance bottlenecks. Worse still, these solutions degrade sharply under variable or weak connectivity, where untimely code delivery can stall execution altogether. We propose HyStream, a hybrid application streaming system that combines predictive remote delivery with local sequence-linearized caching to sustain responsive and robust execution without requiring installation. HyStream addresses three key challenges: (1) maintaining microsecond-level latency comparable to local storage; (2) bridging the semantic gap between stateless remote storage and stateful execution; and (3) mitigating the limitations of purely network-based solutions under degraded connectivity. To achieve this, HyStream integrates three core components: a dual-mode transmission mechanism that decouples synchronous demand-driven requests from asynchronous speculative prefetching; a thread-aware Markov-chain model that captures fine-grained, concurrent access patterns for accurate prediction; and a sequence-linearized cache that persists streamed blocks in predicted execution order to support deterministic fallback behavior. Together, these components transform irregular, latency-sensitive I/O into efficient structured access that masks network variability. Evaluation shows HyStream delivers near-native performance across diverse networks. On mobile devices, it achieves 16–29% better per-page access latency than local UFS3.1, even over variable Wi-Fi connectivity. On desktops, it typically sustains startup overheads below 30% relative to local NVMe. Under variable and degraded network conditions, the sequence-linearized cache increasingly serves execution-critical accesses, rendering application performance largely insensitive to network latency and jitter within intra-city and inter-city deployments. Sheng Yue 0001, Xiang Liu 0017, Yongjian Fu 0004, Jialin Li 0001 |
IEEE Trans. Netw. | 5 |
| 2026 | Toward Communication-Efficient and Data-Free Collaborative Fine-Tuning Between Small and Large Language ModelsabstractWhile large language models (LLMs) exhibit impressive general capabilities, their performance on domainspecific tasks often requires fine-tuning with private data that cannot be shared due to privacy constraints. Directly deploying LLMs on resource-constrained clients for local fine-tuning is impractical due to their significant computation and communication costs. In addition, pre-trained LLMs are valuable intellectual property, and model owners are reluctant to distribute full model weights. To address these challenges, we proposeCoT-LM, a communication-efficient, computation-light, and data-free framework for collaborative fine-tuning between small (SLMs) and large language models (LLMs). InCoT-LM, clients fine-tune lightweight SLMs locally without uploading models or private data. These SLMs provide task-specific feedback to guide server-side LLM enhancement via an efficient communication protocol that exchanges only lightweight synthetic data and feedback. The framework supports both synchronous and asynchronous collaboration and enables mutual enhancement: the LLM improves its task-specific capabilities, while clients benefit from refined synthetic data or distilled knowledge. Extensive experiments demonstrate thatCoT-LMsignificantly boosts natural language understanding (NLU, up to 18.3% for LLMs and 8.0% for SLMs) and natural language generation (NLG, up to 31.7% for LLMs) performance across diverse tasks while preserving data privacy, model intellectual property, and generalization capabilities, achieving significant reductions in computation and communication overhead. Zhenya Ma, Yongheng Deng, Ziqing Qiao, Yongjian Fu 0004, Sheng Yue 0001, Ju Ren 0001 |
IEEE Trans. Netw. | 4 |
| 2025 | WDNN: Weighted Diffractive Neural Network for Physical-layer RF Signal ProcessingabstractDiffractive neural networks (NNs) have garnered attention for directly implementing wireless signal processing at the physical layer. However, they are limited by a constrained weight learning space and activation functions, which restricts their data processing capabilities. To address this, we propose an RF circuit-based weighted diffraction NN (WDNN) that rivals digital NNs in processing ability. We design a weighted asymmetric RF coupler unit that, when stacked into a network, enables diffractive propagation with arbitrary connection weights. Additionally, an activation module is introduced that utilizes RF amplifiers operating in their nonlinear regions. We validate the effectiveness of the proposed WDNN through three tasks: 32-level amplitude modulated (AM) signal decoding, 31-class angle of arrival (AoA) estimation, and 2-class Wi-Fi based fall detection. After training, WDNN achieves the accuracy of 98.5%, 93.7%, and 90.8% in the AM decoding, AoA estimation, and fall detection tasks, respectively; while the diffractive NN SOTA achieves only 21.6%, 16.9%, and 63.3%. We also implement the prototypes of WDNN and SOTA, and real-world experimental results demonstrate that our method achieves an average accuracy improvement of up to 76.85% across various tasks compared to SOTA. Yezhou Wang, Yongjian Fu 0004, Hao Pan 0003, Qinyun Hu, Lili Qiu, Yi-Chao Chen 0001, Guangtao Xue, Ju Ren 0001 |
MobiCom | 2 |
| 2025 | Towards Distance-Adaptive Wireless ChargingabstractWireless charging holds significant promise for IoT devices and transportation networks by facilitating convenient and autonomous power supply. Traditional wireless charging technologies have typically adhered to a singular approach, choosing between near-field coupling or far-field radiation. However, our investigations uncover that each method outperforms the other at specific distances. This insight leads us to integrating the advantages of both to enable rapid wireless charging across any distance within the charging range. For this vision, we poses an intriguing question: "Can we develop a system that supports both near-field and far-field charging simultaneously?" Shuning Wang, Linghui Zhong, Yongjian Fu 0004, Sheng Yue 0001, Ju Ren 0001, Yaoxue Zhang |
MobiSys | 5 |
| 2025 | MODepth: Benchmarking Mobile Multi-frame Monocular Depth Estimation with Optical Image StabilizationabstractThis paper presents MODepth, a multi-frame monocular depth estimation system based on the controlled motion of an optical image stabilization (OIS) module. By actively injecting acoustic signals, we induce regular translational movements of the OIS lens, resulting in controllable camera pose changes and simplifying inter-frame pose estimation. Leveraging multi-frame images captured under OIS-controlled lens movements, we design a high-precision depth estimation network, MODNet, and introduce the principal point offset estimation module and pose estimation modules to fully exploit geometric information across frames. To validate the effectiveness of our approach, we collect a new dataset MODdata with 1100 samples in nearly 220 indoor scenarios and benchmark our model as an OIS-based multi-frame depth estimation method, comparing it to ground truth obtained from a depth sensor and other state-of-the-art monocular depth estimation algorithms. Our method achieves competitive or superior performance compared to fully supervised baselines, reaching an RMSE of 0.439, which outperforms all evaluated methods, demonstrating that self-supervised fine-tuning with OIS-induced parallax is a viable alternative to ground-truth supervision. Code and dataset are available at: https://github.com/liangjindeamo-yuer/MODEPTH Yu Lu 0022, Hao Pan 0003, Dian Ding, Jiatong Ding, Yongjian Fu 0004, Yi-Chao Chen 0001, Ju Ren 0001, Guangtao Xue |
SIGGRAPH Asia | 5 |
| 2025 | UltraPoser: Pushing the Limits of IMU-based Full-Body Pose Estimation with Ultrasound Sensing on Consumer WearablesabstractFigure 1: UltraPoser enables ubiquitous full-body pose estimation by integrating ultrasound sensing and IMU using commodity wearable devices.In addition to measuring IMU data, a smartphone and smartwatch are used to transmit and receive ultrasound signals.The extracted ultrasound features capture motions from joints without any attached devices and offer drift-free range measurements to complement IMU data for more accurate pose estimation. Shuning Wang, Yongjian Fu 0004, Ju Ren 0001, Xinyu Zhang 0003, Akshay Gadre, Ke Sun 0012 |
UIST | 3 |
| 2025 | MoiréComm: Secure Screen-Camera Communication Based on Moiré CryptographyabstractQuick Response (QR) codes have become increasingly popular for screen-camera communication due to their swift readability and widespread smartphone use. Nevertheless, they are vulnerable to privacy invasions from unauthorized photography. Addressing this, we propose a novel Moiré encryption technique-based secure screen-camera communication system, named MoiréComm. The Moiré encryption can enhance security by using distinct spatial frequency patterns for camouflage. The original QR code is revealed as a Moiré pattern only when the camera in a designated position, e.g., directly in front and 30 cm from the screen. From any other positions, only the camouflaged QR code can be seen. Decryption schemes are customized for different scenarios. The multi-frame approach achieves a decryption success of over 98.6% within 13.2 frames in handheld scenarios. Conditional generative adversarial network (cGAN)-based decryption method decodes the Moiré QR code images with a 98.8% success rate in 0.02 s within three frames and is also applicable in handheld scenarios. For fixed screen-camera setups, our fast decryption scheme achieves 99.4% success within two frames, with average 0.4 s latency. Significantly, the decryption rate plunges to 0% for surveillance cameras displaced by 20$^\circ$or more than$\ge$10 cm from the target position, demonstrating MoiréComm's resilience against attacks. Hao Pan 0003, Yongjian Fu 0004, Yu Lu 0022, Feitong Tan, Yi-Chao Chen 0001, Ju Ren 0001 |
IEEE Trans. Dependable Secur. Comput. | 2 |
| 2025 | MagSpy: Revealing User Privacy Leakage via Magnetometer on Mobile DevicesabstractVarious characteristics of mobile applications (apps) and associated in-app services can reveal potentially-sensitive user information; however, privacy concerns have prompted third-party apps to restrict access to data related to mobile app usage. This paper outlines a novel approach to extracting detailed app usage information by analyzing electromagnetic (EM) signals emitted from mobile devices during app-related tasks. The proposed system, MagSpy, recovers user privacy information from magnetometer readings that do not require access permissions. This EM leakage becomes complex when multiple apps are used simultaneously and is subject to interference from geomagnetic signals generated by device movement. To address these challenges, MagSpy employs multiple techniques to extract and identify signals related to app usage. Specifically, the geomagnetic offset signal is canceled using accelerometer and gyroscope sensor data, and a Cascade-LSTM algorithm is used to classify apps and in-app services. MagSpy also uses CWT-based peak detection and a Random Forest classifier to detect PIN inputs. A prototype system was evaluated on over 50 popular mobile apps with 30 devices. Extensive evaluation results demonstrate the efficacy of MagSpy in identifying in-app services (96% accuracy), apps (93.5% accuracy), and extracting PIN input information (96% top-3 accuracy). Yongjian Fu 0004, Lanqing Yang, Hao Pan 0003, Yi-Chao Chen 0001, Guangtao Xue, Ju Ren 0001 |
IEEE Trans. Mob. Comput. | 1 |
| 2025 | MASA: Multimodal Federated Learning Through Modality-Aware and Secure AggregationabstractAs a promising paradigm, federated learning has been applied to multimodal sensing tasks due to its deployment convenience. However, the recent advances in multimodal federated learning emphasize learning a high-quality multimodal model but overlook the model usage requirements of massive unimodal clients. Moreover, the privacy risk in model sharing and client data heterogeneity impact the efficacy of federated learning. In this paper, we propose a novel multimodal federated learning system named MASA. As a departure from existing approaches, MASA simultaneously enhances the model learning efficiency of both multimodal and unimodal clients while ensuring their data privacy. First, we employ a gated cross-modal distillation scheme to achieve performance-aware knowledge transfer across modality-heterogeneous clients. To enhance the system security, MASA integrates a lightweight split-shuffle mechanism to realize the anonymization and encryption of model aggregation. Moreover, to reach personalized collaboration while protecting privacy, MASA features an attention-based spontaneous client clustering mechanism to form client cluster structures securely and distributedly. We evaluate our MASA on four public multimodal datasets for human activity recognition. The results show that our MASA outperforms leading multimodal federated learning methods on the model performance of both multimodal and unimodal clients. Jialin Guo, Yongjian Fu 0004, Zhiwei Zhai, Xinyi Li 0005, Yongheng Deng, Sheng Yue 0001, Hao Pan 0003, Ju Ren 0001 |
IEEE Trans. Mob. Comput. | 2 |
| 2025 | MagicWrite: One-Dimensional Acoustic Tracking-Based Air Writing SystemabstractAir writing technology enhances text input for IoT, VR, and AR devices, offering a spatially flexible alternative to physical keyboards. Addressing the demand for such innovation, this paper presents MagicWrite, a novel system utilizing acoustic-based 1D tracking, which is suitable for mobile devices with existing speaker and microphone infrastructure. Compared to 2D or 3D tracking of the finger, 1D tracking eliminates the need for multiple microphones and/or speakers and is more universally applicable. However, challenges emerge when using 1D tracking for recognizing handwritten letters due to trajectory loss and inter-user writing variability. To address this, we develop a general conversion technique that transforms image-based text datasets (e.g., MNIST) into 1D tracking trajectory data, generating artificial datasets of tracking traces (referred to asTrackMNISTs) to bolster system robustness and scalability. These tracking datasets facilitate the creation of personalized user databases that align with individual writing habits. Combined with a kNN classifier, our proposed MagicWrite ensures high accuracy and robustness in text input recognition while simultaneously reducing computational load and energy consumption. Extensive experiments validate that our proposed MagicWrite achieves exceptional classification accuracy for unseen users and inputs in five languages, marking it as a robust solution for air writing. Hao Pan 0003, Yongjian Fu 0004, Ye Qi, Yi-Chao Chen 0001, Ju Ren 0001 |
IEEE Trans. Mob. Comput. | 2 |
| 2024 | Pushing Wireless Charging from Station to TravelabstractWireless charging has achieved promising progress in recent years. However, the severe bottlenecks are the small charging range and poor flexibility. This paper presents ChargeX to enable smart and long-range wireless charging for small mobile devices. ChargeX incorporates emerging smart metasurface into the magnetic resonance coupling-based wireless charging to extend the charging range and accommodates the mobility of charging device. Unlike previous endeavors in metasurface-assisted wireless charging that focused on simulation, ChargeX makes efforts across software and hardware to meet three crucial requirements for a practical wireless charging system: (i) realize high-freedom and accurate metasurface control under the premise of low loss; (ii) obtain real-time feedback from the receiver and make effective manipulation for transmitted magnetic flux; and (iii) generate a proper AC signal source at the desired frequency band. We developed a prototype of ChargeX, and evaluated its performance through controlled experiments and real-world phone charging. Extensive experiments demonstrate the great potential of ChargeX for long-range and flexible wireless charging with a compact receiver design. Bozhong Yu, Yongjian Fu 0004, Ju Ren 0001, Hao Pan 0003, Jeremy Gummeson, Yaoxue Zhang |
MobiCom | 3 |
| 2024 | Adaptive Metasurface-Based Acoustic Imaging using Joint OptimizationabstractAcoustic imaging is attractive due to its ability to work under occlusion, different lighting conditions, and privacy-sensitive environments. Existing acoustic imaging methods require large transceiver arrays or device movement, which makes it challenging to use in many scenarios. In this paper, we develop a novel acoustic imaging system for low-cost devices with few speakers and microphones without any device movement. To achieve this goal, we leverage a 3D-printed passive acoustic metasurface to significantly enhance the diversity of the measurement data, thereby improving the imaging quality. Specifically, we jointly design the transmission signal, transceivers' beamforming weights, metasurface, and imaging algorithm to minimize the imaging reconstruction error in an end-to-end manner. We further develop a scheme to dynamically adapt the imaging resolution based on the distance to the target. We implement a system prototype. Using extensive experiments, we show that our system yields high-quality images across a wide range of scenarios. Yongjian Fu 0004, Yongzhao Zhang, Yu Lu 0022, Lili Qiu, Yi-Chao Chen 0001, Yezhou Wang, Yijie Li 0002, Ju Ren 0001, Yaoxue Zhang |
MobiSys | 1 |
| 2024 | M3Cam: Extreme Super-resolution via Multi-Modal Optical Flow for Mobile CamerasabstractThe demand for ultra-high-resolution imaging in mobile phone photography is continuously increasing. However, the image resolution of mobile devices is typically constrained by the size of the CMOS sensor. Although deep learning-based super-resolution (SR) techniques have the potential to overcome this limitation, existing SR neural network models require large computational resources, making them unsuitable for real-time SR imaging on current mobile devices. Additionally, cloud-based SR systems pose privacy leakage risks. In this paper, we propose M3Cam, an innovative and lightweight SR imaging system for mobile phones. M3Cam can ensure high-quality 16× SR image (4× in both height and width) visualization with almost negligible latency. In detail, we utilize an optical image stabilization (OIS) module for lens control and introduce a new modality of data, namely gyroscope readings, to achieve high-precision and compact optical flow estimation modules. Building upon this concept, we design a multi-frame-based SR model utilizing the Swin Transformer. Our proposed system can generate a 16× SR image from four captured low-resolution images in real-time, with low computational load, low inference latency, and minimal reliance on runtime RAM. Through extensive experiments, we demonstrate that our proposed multi-modal optical flow model significantly enhances pixel alignment accuracy between multiple frames and delivers outstanding 16× SR imaging results under various shooting scenarios. Code and dataset are available at: https://github.com/liangjindeamo-yuer/M3CAM Yu Lu 0022, Dian Ding, Hao Pan 0003, Yongjian Fu 0004, Feitong Tan, Yi-Chao Chen 0001, Guangtao Xue, Ju Ren 0001 |
SenSys | 4 |
| 2024 | HandPad: Make Your Hand an On-the-go Writing Pad via Human CapacitanceabstractThe convenient text input system is a pain point for devices such as AR glasses, and it is difficult for existing solutions to balance portability and efficiency. This paper introduces HandPad, the system that turns the hand into an on-the-go touchscreen, which realizes interaction on the hand via human capacitance. HandPad achieves keystroke and handwriting inputs for letters, numbers, and Chinese characters, reducing the dependency on capacitive or pressure sensor arrays. Specifically, the system verifies the feasibility of touch point localization on the hand using the human capacitance model and proposes a handwriting recognition system based on Bi-LSTM and ResNet. The transfer learning-based system only needs a small amount of training data to build a handwriting recognition model for the target user. Experiments in real environments verify the feasibility of HandPad for keystroke (accuracy of 100%) and handwriting recognition for letters (accuracy of 99.1%), numbers (accuracy of 97.6%) and Chinese characters (accuracy of 97.9%). Yu Lu 0022, Dian Ding, Hao Pan 0003, Yijie Li 0002, Juntao Zhou, Yongjian Fu 0004, Yongzhao Zhang, Yi-Chao Chen 0001, Guangtao Xue |
UIST | 6 |
| 2024 | UltraSR: Silent Speech Reconstruction via Acoustic SensingabstractSilent Speech Interfaces (SSI) have been developed to convert silent articulatory gestures into speech, aiding communication in public spaces and assisting individuals with aphasia. Previous SSIs, which rely on wearable devices or cameras, often pose issues like prolonged contact or privacy risks. Recent advancements in acoustic sensing present new opportunities for gesture sensing, but they typically focus on content classification rather than reconstructing audible speech. This results in the loss of crucial speech characteristics such as rate, intonation, and emotion.In this paper, we propose UltraSR, a novel sensing system designed for accurate audible speech reconstruction by analyzing the disturbance of tiny articulatory gestures on reflected ultrasound signals. UltraSR employs a multi-scale feature extraction scheme to aggregate information from multiple views and introduces a new model that maps ultrasound to speech signals, enabling the reconstruction of audible speech from silent gestures.Instead of the laborious collection of massive training data, UltraSR constructs an inverse task to generate virtual gestures from widely available audio (e.g., phone calls) for efficient model training. Additionally, it incorporates a finetuning mechanism using unlabeled data for user adaptation.We implemented UltraSR on a portable smartphone and evaluated it in various environments. Results show that UltraSR can achieve a Character Error Rate (CER) as low as 5.22% and reduce the CER from 80.13% to 6.31% for new users with only 1 hour of ultrasound data, outperforming state-of-the-art acoustic-based approaches while preserving rich speech information. Yongjian Fu 0004, Shuning Wang, Linghui Zhong, Ju Ren 0001, Yaoxue Zhang |
IEEE Trans. Mob. Comput. | 1 |
| 2022 | SVoice: Enabling Voice Communication in Silence via Acoustic Sensing on Commodity DevicesabstractSilent Speech Interface (SSI) has been proposed as a means of reconstructing audible speech from silent articulatory gestures for covert voice communication in public and voice assistance for the aphasic. Prior arts of SSI, either relying on wearable devices or cameras, may lead to extended contact requirements or privacy leakage risks. The recent advances in acoustic sensing have brought new opportunities for sensing gestures, but their original intention is to infer speech content for classification instead of audible speech reconstruction, resulting in the loss of some important speech information (e.g., speech rate, intonation, and emotion). In this paper, we propose, the first system that supports accurate audible speech reconstruction by analyzing the disturbance of tiny articulatory gestures on the reflected ultrasound signal. The design of introduces a new model that provides the unique mapping relationship between ultrasound and speech signals, so that the audible speech can be successfully reconstructed from the silent speech. However, establishing the mapping relationship depends on plenty of training data. Instead of the time-consuming collection of massive amounts of data for training, we construct an inverse task that constitutes a dual form with the original task to generate virtual gestures from widely available audio (e.g., phone calls) for facilitating model training. Furthermore, we introduce a fine-tuning mechanism using unlabeled data for user adaptation. We implement using a portable smartphone and evaluate it in various environments. The evaluation results show that can reconstruct speech with a (Character Error Rate) CER as low as 7.62%, and decrease the CER from 82.77% to 9.42% on new users with only 1 hour of ultrasound signals provided, which outperforms state-of-the-art acoustic-based approaches while preserving rich speech information. Yongjian Fu 0004, Shuning Wang, Linghui Zhong, Ju Ren 0001, Yaoxue Zhang |
SenSys | 1 |