Yu Lu 0022

dblp:09/2321-22 · DBLP profile ↗
← Back
17ranked-venue papers
7as first author
17since 2021 · last 2026
0000-0002-9024-3692ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 9 · 4 first-author · 9 since 2021Human-computer interaction and ubiquitous computing · 3 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-author · 2 since 2021Security and privacy · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Aucom: Extreme Compression for Real-Time Edge-to-Server Universal Audio Streaming
abstract
Real-time audio streaming transmission and processing play a crucial role in time-sensitive applications such as food delivery services and ride-hailing platforms, where rapid response is essential. However, existing server-based audio streaming architectures struggle to handle the high concurrency of massive mobile devices efficiently. Traditional compression methods like MP3 and AAC offer limited compression ratios, while deep learning-based approaches often fail to meet the real-time transmission demands of edge computing environments. In this paper, we propose a novel edge-to-server audio streaming architecture that leverages Mel filter bank spectral features to achieve ultra-high compression efficiency. Our system integrates audio denoising, Mel feature extraction, and quantization-based compression at the edge, effectively suppressing environmental and device-induced noise while achieving an extreme compression ratio of 0.39% relative to the original uncompressed audio. Compared to conventional methods like MP3, our approach further reduces the file size by 96.1%. The decompressed Mel features remain task-independent, enabling seamless support for various general-purpose audio processing tasks in the server. We evaluate our system across three key audio tasks: speech recognition, speech emotion recognition, and audio classification. Extensive experiments on five different mobile devices demonstrate a 93.10% reduction in transmission latency at 1 Mbps bandwidth compared to 64 kbps MP3 audio, while maintaining task performance within a 5% deviation from state-of-the-art (SOTA) models across six mainstream audio datasets. These results highlight the efficiency, robustness, and scalability of our approach for real-time edge-to-server audio processing.
Yu Lu 0022, Dian Ding, Yijie Li 0002, Longyuan Ge, Juntao Zhou, Yongzhao Zhang, Yi-Chao Chen 0001, Jiannong Cao 0001, Guangtao Xue
IEEE Trans. Mob. Comput.1
2025 STELLAR: Pacemaker Recognition Using 12-Lead ECG and Spatio-Temporal Harmonic Mechanism
abstract
As cardiovascular diseases and arrhythmias rise globally, pacemakers have become a critical therapeutic option for managing cardiac rhythm disorders. Accurate identification of pacemaker implantation sites is essential for personalized pacing therapy and optimal clinical outcomes. While 12-lead electrocardiogram (ECG) signals provide a non-invasive means to infer implantation locations, they are susceptible to noise and morphological variability, posing challenges for high-accuracy localization. To advance data-driven solutions in this domain, we present PILDE, the first publicly available dataset specifically designed for pacemaker implantation site identification, comprising 12-lead ECG recordings from 733 patients across four distinct implantation locations. Based on this dataset, we propose STELLAR, a novel deep learning framework that integrates a Spatio-Temporal Lead-Harmonic Mechanism to model both the temporal dynamics of ECG waveforms and the spatial coherence across leads. Extensive experiments demonstrate that STELLAR outperforms conventional deep models-including CNN, LSTM, and Transformer baselines-on both the PILDE and PTB-XL datasets. Specifically, STELLAR achieves an average accuracy improvement of 10.45 % on PILDE and 14.19 % on PTB-XL, with significant gains in sensitivity and F1-score for minority classes. These results highlight the robustness and precision of STELLAR in automating implantation site identification, offering a promising tool for pre-procedural planning and clinical decision support. The source code and dataset access information will be made publicly available.
Han Zhang 0053, Zeyuan Ding, Leping Yang, Yu Lu 0022, Jiatong Ding, Dian Ding, Yiding Qi, Ruogu Li, Guanghui Gao, Yi-Chao Chen 0001, Guangtao Xue
BIBM4
2025 A Transform-Domain Approach with Symmetric and Edge Constraints for MRI Super-Resolution
abstract
Magnetic resonance imaging (MRI) provides highquality soft tissue contrast images and is crucial in medical diagnosis. However, systems face trade-offs between image resolution and scan time. Low-resolution MRI scans reduce scan time and patient burden but lose critical details needed for accurate diagnosis. To address this problem, super-resolution techniques have been developed to improve the clarity of lowresolution input images. Single-image super-resolution (SISR), which minimizes patient scanning time, has gradually become a research focus, but existing methods often struggle to balance the reconstruction of low-frequency structural information and high-frequency details. In this paper, we propose a novel superresolution up-sampling pipeline that enhances both the highfrequency and low-frequency components of magnetic resonance imaging. In addition, we introduce an enhanced loss function that includes symmetry and edge constraints to preserve critical structural details for improved diagnostic accuracy. The extensive experiments across multiple datasets validate the effectiveness of our SISR model. Source code will be made publicly available.
Han Zhang 0053, Yu Lu 0022, Dian Ding, Mengying Zhu, Shengyun He, Yi-Chao Chen 0001, Ruokun Li, Shikui Tu, Guangtao Xue
BIBM2
2025 M2SILENT: Enabling Multi-user Silent Speech Interactions via Multi-directional Speakers in Shared Spaces
abstract
We introduce M 2 Silent, which enables multi-user silent speech interactions in shared spaces using multi-directional speakers.Ensuring privacy during interactions with voice-controlled systems presents significant challenges, particularly in environments with multiple individuals, such as libraries, offices, or vehicles.M 2 Silent addresses this by allowing users to communicate silently, without producing audible speech, using acoustic sensing integrated into directional speakers.We leverage FMCW signals as audio carriers, simultaneously playing audio and sensing the user's silent speech.
Juntao Zhou, Dian Ding, Yijie Li 0002, Yu Lu 0022, Yida Wang 0007, Yongzhao Zhang, Yi-Chao Chen 0001, Guangtao Xue
CHI4
2025 AMSER: Accelerate Mobile Speech Emotion Recognition with Signal Compression
abstract
Speech-based interaction systems are widely used in mobile devices like smartphones. With advances in deep neural networks, tasks such as speech emotion recognition (SER) enhance these systems’ user-friendliness. However, deploying SER models on mobile devices is challenging due to their complexity and computational demands. While pruning can reduce complexity, it often compromises accuracy, and hardware accelerators like FPGAs are difficult to integrate into mobile devices. This paper proposes AMSER, a real-time speech emotion recognition framework using signal compression and task offloading. AMSER utilizes logarithmic Mel-filter bank coefficients (Fbank) and singular value decomposition (SVD) for feature extraction and compression. The compressed signal is only 6.25% of the original size, achieving 2.24x faster transfer rates and 55.35% energy savings compared to raw audio transmission. Despite the compression, the features preserve key audio information for text and emotion recognition, performed server-side. Experiments show a WER of 4.68% (Librispeech), 10.69% (CommonVoice), and 69.83% emotion recognition accuracy (IEMOCAP).
Yu Lu 0022, Dian Ding, Han Zhang 0053, Lanqing Yang, Yi-Chao Chen 0001, Guangtao Xue
ICASSP1
2025 High-resolution mmWave Imaging using Metasurface and Diffusion
Yida Wang 0007, Yu Lu 0022, Yifei Shen 0004, Lili Qiu, Zeyuan Lai, Yi-Chao Chen 0001, Hao Pan 0003, Juntao Zhou, Dian Ding, Guangtao Xue, Qian Zhang 0001
MobiSys2
2025 MODepth: Benchmarking Mobile Multi-frame Monocular Depth Estimation with Optical Image Stabilization
abstract
This paper presents MODepth, a multi-frame monocular depth estimation system based on the controlled motion of an optical image stabilization (OIS) module. By actively injecting acoustic signals, we induce regular translational movements of the OIS lens, resulting in controllable camera pose changes and simplifying inter-frame pose estimation. Leveraging multi-frame images captured under OIS-controlled lens movements, we design a high-precision depth estimation network, MODNet, and introduce the principal point offset estimation module and pose estimation modules to fully exploit geometric information across frames. To validate the effectiveness of our approach, we collect a new dataset MODdata with 1100 samples in nearly 220 indoor scenarios and benchmark our model as an OIS-based multi-frame depth estimation method, comparing it to ground truth obtained from a depth sensor and other state-of-the-art monocular depth estimation algorithms. Our method achieves competitive or superior performance compared to fully supervised baselines, reaching an RMSE of 0.439, which outperforms all evaluated methods, demonstrating that self-supervised fine-tuning with OIS-induced parallax is a viable alternative to ground-truth supervision. Code and dataset are available at: https://github.com/liangjindeamo-yuer/MODEPTH
Yu Lu 0022, Hao Pan 0003, Dian Ding, Jiatong Ding, Yongjian Fu 0004, Yi-Chao Chen 0001, Ju Ren 0001, Guangtao Xue
SIGGRAPH Asia1
2025 Amser+: Accelerating Mobile Speech Emotion Recognition in IoT Environments With Mel Feature Compression
abstract
Speech-based interaction systems are widely used in mobile devices like smartphones. With advances in deep neural networks, tasks such as speech emotion recognition (SER) enhance these systems user-friendliness. However, deploying SER models on mobile devices is challenging due to their complexity and computational demands. While pruning can reduce complexity, it often compromises accuracy, and hardware accelerators like FPGAs are difficult to integrate into mobile devices. This paper proposes Amser+, a real-time speech emotion recognition framework using signal compression and task offloading. Amser+utilizes logarithmic Mel-filter bank coefficients (Fbank) and singular value decomposition (SVD) for feature extraction and compression. The compressed signal is only 6.25% of the original size, achieving 2.24× faster transfer rates and 55.35% energy savings compared to raw audio transmission. Despite the compression, the features preserve key audio information for text and emotion recognition, performed server-side. Experiments show a WER of 4.68% (Librispeech), 10.69% (CommonVoice), and 72.85% emotion recognition accuracy (IEMOCAP).
Yu Lu 0022, Dian Ding, Yijie Li 0002, Yongzhao Zhang, Lanqing Yang, Yi-Chao Chen 0001, Guangtao Xue
IEEE Internet Things J.1
2025 MoiréComm: Secure Screen-Camera Communication Based on Moiré Cryptography
abstract
Quick Response (QR) codes have become increasingly popular for screen-camera communication due to their swift readability and widespread smartphone use. Nevertheless, they are vulnerable to privacy invasions from unauthorized photography. Addressing this, we propose a novel Moiré encryption technique-based secure screen-camera communication system, named MoiréComm. The Moiré encryption can enhance security by using distinct spatial frequency patterns for camouflage. The original QR code is revealed as a Moiré pattern only when the camera in a designated position, e.g., directly in front and 30 cm from the screen. From any other positions, only the camouflaged QR code can be seen. Decryption schemes are customized for different scenarios. The multi-frame approach achieves a decryption success of over 98.6% within 13.2 frames in handheld scenarios. Conditional generative adversarial network (cGAN)-based decryption method decodes the Moiré QR code images with a 98.8% success rate in 0.02 s within three frames and is also applicable in handheld scenarios. For fixed screen-camera setups, our fast decryption scheme achieves 99.4% success within two frames, with average 0.4 s latency. Significantly, the decryption rate plunges to 0% for surveillance cameras displaced by 20$^\circ$or more than$\ge$10 cm from the target position, demonstrating MoiréComm's resilience against attacks.
Hao Pan 0003, Yongjian Fu 0004, Yu Lu 0022, Feitong Tan, Yi-Chao Chen 0001, Ju Ren 0001
IEEE Trans. Dependable Secur. Comput.3
2025 TouchHBC: Touch-Based Human Body Communication via Leakage Current
abstract
Wearable devices, including smartwatches, are increasingly popular among consumers due to their user-friendly services. However, transmitting sensitive data like social media messages and payment QR codes via commonly used low-power Bluetooth exposes users to privacy breaches and financial losses. This study introducesTouchHBC, a secure and reliable communication scheme leveraging a smartwatch's built-in electrodes. This system establishes a touch-based human communication system utilizing a laptop's leakage current. As the transmitting device, the laptop modulates this current via the CPU. Simultaneously, the smartwatch, equipped with built-in electrodes, captures the current traversing the human body and decodes it. The modulation and decoding processes involve techniques such as amplitude modulation, variational mode decomposition, channel estimation, and retransmission mechanisms.TouchHBCfacilitates communication between laptops and smartwatches. Real-world tests demonstrate that our prototype achieves a throughput of$19.83bps$. Moreover,TouchHBCoffers the potential for enhanced interaction, including improved gaming experiences through vibration feedback and secure touch login for smartwatch applications by synchronizing with a laptop. Furthermore, the system can be integrated with high-throughput communication protocols such as Bluetooth, enhancing its scalability while maintaining a strong foundation of security.
Dian Ding, Hao Pan 0003, Yongzhao Zhang, Yijie Li 0002, Yu Lu 0022, Yi-Chao Chen 0001, Guangtao Xue
IEEE Trans. Mob. Comput.5
2024 CarbonNet: Enterprise-Level Carbon Emission Prediction with Large-Scale Datasets
Jinghua Tang, Lanqing Yang, Yuqiao Pei, Dian Ding, Yu Lu 0022, Guangtao Xue
ICIC (12)7
2024 DASIV: Directional Acoustic Sensing based Intelligent Vehicle Interaction System
abstract
With the increase in motor vehicles, more convenient and accurate interactions are expected while retaining a high standard of safe driving. However, complex and dynamic vehicle environments challenge sensing tasks such as breathing monitor and hand gesture recognition. In this paper, we propose DASIV, which utilizes the highly directional nature of ultrasonic signals to achieve fine-grained directional acoustic sensing in vehicle environments. Due to air nonlinearity, the system enables synchronized directional acoustic communication to transmit information (e.g., navigation) to the driver without affecting other passengers. By optimizing the frequency of the Frequency Modulated Continuous Wave (FMCW) signals, DASIV avoids mutual interference between the sensing and communication signals and achieves breathing detection and hand gesture recognition for the driver. Specifically, the system extracts breathing-induced weak thoracic bullying through the signal phase, captures and analyses breathing patterns using bandpass and Gaussian filters, and develops a breathing model. Then, the system defines 10 interaction hand gestures to meet daily interaction needs, uses spectral features to mine complex and fast hand movement features, and proposes a hand gesture recognition model. Extensive experiments in real environments show that DASIV achieves high-precision breathing monitor (Pearson correlation coefficient of 0.89) and hand gesture recognition (Precision of 91.7%).
Dinghua Zhao, Juntao Zhou, Dian Ding, Yu Lu 0022, Yijie Li 0002, Yi-Chao Chen 0001, Guangtao Xue
IPCCC4
2024 Adaptive Metasurface-Based Acoustic Imaging using Joint Optimization
abstract
Acoustic imaging is attractive due to its ability to work under occlusion, different lighting conditions, and privacy-sensitive environments. Existing acoustic imaging methods require large transceiver arrays or device movement, which makes it challenging to use in many scenarios. In this paper, we develop a novel acoustic imaging system for low-cost devices with few speakers and microphones without any device movement. To achieve this goal, we leverage a 3D-printed passive acoustic metasurface to significantly enhance the diversity of the measurement data, thereby improving the imaging quality. Specifically, we jointly design the transmission signal, transceivers' beamforming weights, metasurface, and imaging algorithm to minimize the imaging reconstruction error in an end-to-end manner. We further develop a scheme to dynamically adapt the imaging resolution based on the distance to the target. We implement a system prototype. Using extensive experiments, we show that our system yields high-quality images across a wide range of scenarios.
Yongjian Fu 0004, Yongzhao Zhang, Yu Lu 0022, Lili Qiu, Yi-Chao Chen 0001, Yezhou Wang, Yijie Li 0002, Ju Ren 0001, Yaoxue Zhang
MobiSys3
2024 M3Cam: Extreme Super-resolution via Multi-Modal Optical Flow for Mobile Cameras
abstract
The demand for ultra-high-resolution imaging in mobile phone photography is continuously increasing. However, the image resolution of mobile devices is typically constrained by the size of the CMOS sensor. Although deep learning-based super-resolution (SR) techniques have the potential to overcome this limitation, existing SR neural network models require large computational resources, making them unsuitable for real-time SR imaging on current mobile devices. Additionally, cloud-based SR systems pose privacy leakage risks. In this paper, we propose M3Cam, an innovative and lightweight SR imaging system for mobile phones. M3Cam can ensure high-quality 16× SR image (4× in both height and width) visualization with almost negligible latency. In detail, we utilize an optical image stabilization (OIS) module for lens control and introduce a new modality of data, namely gyroscope readings, to achieve high-precision and compact optical flow estimation modules. Building upon this concept, we design a multi-frame-based SR model utilizing the Swin Transformer. Our proposed system can generate a 16× SR image from four captured low-resolution images in real-time, with low computational load, low inference latency, and minimal reliance on runtime RAM. Through extensive experiments, we demonstrate that our proposed multi-modal optical flow model significantly enhances pixel alignment accuracy between multiple frames and delivers outstanding 16× SR imaging results under various shooting scenarios. Code and dataset are available at: https://github.com/liangjindeamo-yuer/M3CAM
Yu Lu 0022, Dian Ding, Hao Pan 0003, Yongjian Fu 0004, Feitong Tan, Yi-Chao Chen 0001, Guangtao Xue, Ju Ren 0001
SenSys1
2024 HandPad: Make Your Hand an On-the-go Writing Pad via Human Capacitance
abstract
The convenient text input system is a pain point for devices such as AR glasses, and it is difficult for existing solutions to balance portability and efficiency. This paper introduces HandPad, the system that turns the hand into an on-the-go touchscreen, which realizes interaction on the hand via human capacitance. HandPad achieves keystroke and handwriting inputs for letters, numbers, and Chinese characters, reducing the dependency on capacitive or pressure sensor arrays. Specifically, the system verifies the feasibility of touch point localization on the hand using the human capacitance model and proposes a handwriting recognition system based on Bi-LSTM and ResNet. The transfer learning-based system only needs a small amount of training data to build a handwriting recognition model for the target user. Experiments in real environments verify the feasibility of HandPad for keystroke (accuracy of 100%) and handwriting recognition for letters (accuracy of 99.1%), numbers (accuracy of 97.6%) and Chinese characters (accuracy of 97.9%).
Yu Lu 0022, Dian Ding, Hao Pan 0003, Yijie Li 0002, Juntao Zhou, Yongjian Fu 0004, Yongzhao Zhang, Yi-Chao Chen 0001, Guangtao Xue
UIST1
2023 Effectively Learning Moiré QR Code Decryption from Simulated Data
Yu Lu 0022, Hao Pan 0003, Feitong Tan, Yi-Chao Chen 0001, Jiadi Yu, Jinghai He, Guangtao Xue
INFOCOM1
2023 Addressing Practical Challenges in Acoustic Sensing To Enable Fast Motion Tracking
abstract
Motivated by many potential applications that could be enabled by acoustic motion tracking, in this paper we systematically examine the factors that limit the accuracy of acoustic tracking in practical scenarios. We identify three main challenges: (i) high mobility, (ii) low SNR, and (iii) hardware frequency response. We further show that the last two issues may exacerbate the performance issue under high mobility. We develop effective approaches to address the issues. In particular, to address high mobility, we tackle phase wrap-around using the derivative of the phase; we further estimate the Doppler shift under diverse scenarios and compensate the Doppler in channel impulse response (CIR). To address low SNR, we use a novel approach to estimate the phase shift between consecutive time intervals to effectively support time-domain beamforming and increase SNR. To tackle the uneven frequency response, we show that it is important to estimate and compensate the phase as well as the amplitude of the frequency response. Our extensive evaluation shows that each of our techniques is effective and putting them together significantly enhances the accuracy of acoustic motion tracking in general scenarios.
Yongzhao Zhang, Hao Pan 0003, Yi-Chao Chen 0001, Lili Qiu, Yu Lu 0022, Guangtao Xue, Jiadi Yu, Feng Lyu 0001
IPSN5