Dian Ding

dblp:271/9812 · DBLP profile ↗
← Back
35ranked-venue papers
5as first author
34since 2021 · last 2026
0000-0002-2190-0919ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 18 · 5 first-author · 17 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 5 since 2021Systems, architecture and hardware · 4 · 4 since 2021Human-computer interaction and ubiquitous computing · 3 · 3 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 ShadowClone: Accelerating Cross-Shard Transactions via Shadow Accounts
Jiahao Qi, Dian Ding, Feilong Lin, Jie Li 0002, Shengyun Liu, Guangtao Xue, Jiannong Cao 0001
ICDCS2
2026 BIND: Enabling Continuous Transaction Processing During Account Migration in Sharded Blockchains
abstract
Account migration in sharded blockchains presents a critical trade-off between optimization effectiveness and system availability. While dynamically reallocating accounts across shards can significantly reduce cross-shard transaction overhead, existing migration mechanisms cause service disruptions that intensify as state data volumes grow. To address this challenge, we propose BIND, a batch-wise account migration protocol that eliminates service interruptions by enabling continuous transaction processing throughout migration. BIND introduces a dual transaction pool architecture that isolates transactions involving migrating accounts while allowing non-migrating accounts to operate uninterrupted. To optimize migration efficiency, we design a reverse greedy heuristic algorithm that partitions accounts into batches based on community cohesion, maximizing intra-batch connectivity to front-load cross-shard communication reduction. We evaluate BIND using real Ethereum transactions, demonstrating superior performance over existing mechanisms. BIND achieves 12% higher overall throughput, reduces migration time to 23.6%-39.3% of the one-shot baseline (across 1-10Gbps bandwidth), and lowers cross-shard transaction rates by 24.1% compared to random batching. These results confirm BIND as a practical solution for large-scale, non-disruptive account migration in production sharded blockchains.
Jiahao Qi, Dian Ding, Jie Li 0002, Jiannong Cao 0001, Yi-Chao Chen 0001, Guangtao Xue, Shengyun Liu
WWW2
2026 Exploiting Cyber Threat Intelligence for Indirect Attacks Against Serverless Infrastructures
abstract
Cyber Threat Intelligence (CTI) and serverless computing are two emerging technologies that have significantly impacted their respective domains in recent years. However, their interaction remains surprisingly underexplored. In this work, through in-depth semi-structured interviews with cybersecurity experts, we identify the trust issues within the CTI ecosystem that can be exploited to introduce fake CTI manipulation, enabling indirect attacks against entities with dynamic IP allocation, such as those in serverless computing. Furthermore, these attacks can be amplified by commercial CTI platforms due to their widespread adoption and sharing mechanisms. Based on these insights, we propose Ares, a novel attack strategy that leverages fake CTI manipulation to enable large-scale, stealthy indirect denial-of-service attacks against serverless infrastructures. We demonstrate the feasibility and impact of Ares through extensive evaluations in a controlled experimental environment. Our results show that Ares can rapidly and widely disseminate fake CTI within the CTI ecosystem, leading to an overall average reject rate of 23.03% and a high reject rate of up to 45.42% when accessing top websites in certain industries, while maintaining a low detection rate across state-of-the-art serverless security systems. These findings underscore the urgent need for more frequent communication and collaboration among CTI platforms and related stakeholders to develop a more robust trustworthiness model across the ecosystem.
Baojin Wang, Yongzhao Zhang, Xiong Li 0002, Jie Yang 0003, Ting Chen 0002, Xiaosong Zhang 0001, Dian Ding, Yi-Chao Chen 0001
IEEE Trans. Inf. Forensics Secur.9
2026 Sniffing the Application Usage Information With the Leakage Current of Laptops
abstract
Smart devices are proliferating in every aspect of our lives, providing convenience but also exposing us to the risk of information leakage at any moment. Attackers can monitor the user and infer private information such as personality and preferences by stealing the behavioral information. In this paper, we investigated the potential threat of information stealing via the leakage current of laptops and electrodes in wearable devices (e.g., smart watches and bracelets). Specifically, the leakage current in the laptop adapter can flow from the metal casing into the human body and be collected by electrodes in wearable devices when the user is using a laptop with a metal casing (e.g., MacBook). We verified the correlation between leakage current and the working states of the laptop, where different operations corresponding to different CPU instructions can generate different leakage currents. Based on this, we proposeLeakThief, a system that consists of three components: leakage current detection, application operation detection, and application recognition. The experiments in a real-world environment demonstrated that the proposed system can recognize 25 common applications with high accuracy, including launching-based (96.4%) and in-application operation-based recognition (81.2%).
Dian Ding, Yijie Li 0002, Yongzhao Zhang, Yi-Chao Chen 0001, Xiaoyu Ji 0001, Guangtao Xue
IEEE Trans. Mob. Comput.1
2026 Aucom: Extreme Compression for Real-Time Edge-to-Server Universal Audio Streaming
abstract
Real-time audio streaming transmission and processing play a crucial role in time-sensitive applications such as food delivery services and ride-hailing platforms, where rapid response is essential. However, existing server-based audio streaming architectures struggle to handle the high concurrency of massive mobile devices efficiently. Traditional compression methods like MP3 and AAC offer limited compression ratios, while deep learning-based approaches often fail to meet the real-time transmission demands of edge computing environments. In this paper, we propose a novel edge-to-server audio streaming architecture that leverages Mel filter bank spectral features to achieve ultra-high compression efficiency. Our system integrates audio denoising, Mel feature extraction, and quantization-based compression at the edge, effectively suppressing environmental and device-induced noise while achieving an extreme compression ratio of 0.39% relative to the original uncompressed audio. Compared to conventional methods like MP3, our approach further reduces the file size by 96.1%. The decompressed Mel features remain task-independent, enabling seamless support for various general-purpose audio processing tasks in the server. We evaluate our system across three key audio tasks: speech recognition, speech emotion recognition, and audio classification. Extensive experiments on five different mobile devices demonstrate a 93.10% reduction in transmission latency at 1 Mbps bandwidth compared to 64 kbps MP3 audio, while maintaining task performance within a 5% deviation from state-of-the-art (SOTA) models across six mainstream audio datasets. These results highlight the efficiency, robustness, and scalability of our approach for real-time edge-to-server audio processing.
Yu Lu 0022, Dian Ding, Yijie Li 0002, Longyuan Ge, Juntao Zhou, Yongzhao Zhang, Yi-Chao Chen 0001, Jiannong Cao 0001, Guangtao Xue
IEEE Trans. Mob. Comput.3
2026 MagPrint++: Continuous User Fingerprinting on Mobile Devices Using Electromagnetic Signals
abstract
Understanding the nature of user-device interactions (e.g., who is using the device and what he/she is doing with it) is critical for many applications including time management, user profiles, and privacy protection. However, in scenarios where mobile devices are shared among family members or multiple employees in a company, conventional account-based statistics are not meaningful. This poses an even bigger problem when dealing with sensitive data. Moreover, fingerprint readers and front-facing cameras were not designed to continuously identify users. In this study, we developedMagPrint++, a novel approach to fingerprint users based on unique patterns in the electromagnetic (EM) signals associated with the specific use patterns of users. Initial experiments showed that time-varying EM patterns are unique to individual users. They are also temporally and spatially consistent, which makes them suitable for fingerprinting.MagPrint++has a number of advantages over existing schemes: i) Non-intrusive fingerprinting, ii) implementation both on COTS mobile phones and a small and easy-to-deploy device, and iii) high accuracy thanks to the proposed classification algorithm. In experiments involving 30 users,MagPrint++achieves$94.3\%$accuracy in classifying users from these traces, which represents a$10.9\%$improvement over the state-of-the-art classification method.
Lanqing Yang, Xinqi Chen, Hao Pan 0003, Yi-Chao Chen 0001, Guangtao Xue, Zechen Li 0005, Yiheng Bian, Dian Ding, Linghe Kong, Jiadi Yu, Feng Lyu 0001, Minglu Li 0001, Ziyu Shen, Bo Zhang 0004
IEEE Trans. Mob. Comput.8
2026 LLM4Load-Turbo: A Prompt-Driven LLM Framework With Knowledge Distillation for Efficient Multi-Scale Workload Prediction
abstract
Accurate workload prediction is essential to ensure application Quality of Service (QoS), cost efficiency, and compliance with Service Level Agreements (SLAs) during cloud-based deployment. However, existing methods struggle to achieve accurate forecasts across multiple temporal scales and often fail to generalize well with limited historical data. To tackle these problems, we propose LLM4Load-Turbo, a Prompt-Driven LLM Framework with Knowledge Distillation for Efficient Multi-Scale Workload Prediction. Firstly, we design a structured prompt with dataset introduction, task description, and workload features characterization to extract multi-scale features. Secondly, we introduce a cross-modality alignment mechanism combined with label embedding to further enhance predictive accuracy and generalization, effectively mitigating the container cold-start problem. Thirdly, we propose a two-level knowledge distillation strategy, enabling LLM4Load-Turbo to maintain high accuracy while substantially reducing inference latency, memory footprint, and computational cost. Specifically, our framework achieves up to an 89.27% improvement in inference speed and a 93.30% reduction in model parameter scale compared to state-of-the-art baselines. Extensive experiments on four real cloud workload datasets validate the effectiveness of our framework. For multi-scale workload prediction, LLM4Load-Turbo improves up to 50.44%. For container cold-start scenarios, LLM4Load-Turbo improves up to 93.54%. These results demonstrate the potential of LLM4Load-Turbo to enable dynamic and efficient resource management in modern cloud systems.
Zeyuan Ding, Dian Ding, Jiannong Cao 0001, Yiming Zhang 0003, Guangtao Xue
IEEE Trans. Serv. Comput.2
2025 STELLAR: Pacemaker Recognition Using 12-Lead ECG and Spatio-Temporal Harmonic Mechanism
abstract
As cardiovascular diseases and arrhythmias rise globally, pacemakers have become a critical therapeutic option for managing cardiac rhythm disorders. Accurate identification of pacemaker implantation sites is essential for personalized pacing therapy and optimal clinical outcomes. While 12-lead electrocardiogram (ECG) signals provide a non-invasive means to infer implantation locations, they are susceptible to noise and morphological variability, posing challenges for high-accuracy localization. To advance data-driven solutions in this domain, we present PILDE, the first publicly available dataset specifically designed for pacemaker implantation site identification, comprising 12-lead ECG recordings from 733 patients across four distinct implantation locations. Based on this dataset, we propose STELLAR, a novel deep learning framework that integrates a Spatio-Temporal Lead-Harmonic Mechanism to model both the temporal dynamics of ECG waveforms and the spatial coherence across leads. Extensive experiments demonstrate that STELLAR outperforms conventional deep models-including CNN, LSTM, and Transformer baselines-on both the PILDE and PTB-XL datasets. Specifically, STELLAR achieves an average accuracy improvement of 10.45 % on PILDE and 14.19 % on PTB-XL, with significant gains in sensitivity and F1-score for minority classes. These results highlight the robustness and precision of STELLAR in automating implantation site identification, offering a promising tool for pre-procedural planning and clinical decision support. The source code and dataset access information will be made publicly available.
Han Zhang 0053, Zeyuan Ding, Leping Yang, Yu Lu 0022, Jiatong Ding, Dian Ding, Yiding Qi, Ruogu Li, Guanghui Gao, Yi-Chao Chen 0001, Guangtao Xue
BIBM6
2025 A Transform-Domain Approach with Symmetric and Edge Constraints for MRI Super-Resolution
abstract
Magnetic resonance imaging (MRI) provides highquality soft tissue contrast images and is crucial in medical diagnosis. However, systems face trade-offs between image resolution and scan time. Low-resolution MRI scans reduce scan time and patient burden but lose critical details needed for accurate diagnosis. To address this problem, super-resolution techniques have been developed to improve the clarity of lowresolution input images. Single-image super-resolution (SISR), which minimizes patient scanning time, has gradually become a research focus, but existing methods often struggle to balance the reconstruction of low-frequency structural information and high-frequency details. In this paper, we propose a novel superresolution up-sampling pipeline that enhances both the highfrequency and low-frequency components of magnetic resonance imaging. In addition, we introduce an enhanced loss function that includes symmetry and edge constraints to preserve critical structural details for improved diagnostic accuracy. The extensive experiments across multiple datasets validate the effectiveness of our SISR model. Source code will be made publicly available.
Han Zhang 0053, Yu Lu 0022, Dian Ding, Mengying Zhu, Shengyun He, Yi-Chao Chen 0001, Ruokun Li, Shikui Tu, Guangtao Xue
BIBM4
2025 NLCTCN: A Non-Local Temporal Convolutional Framework for Spatiotemporal Modeling in Multichannel EEG
abstract
Electroencephalography (EEG) analysis plays a crit-ical role in applications such as brain-computer interfaces, epilepsy detection, and cognitive state recognition. However, EEG data are often limited in volume due to high acquisition costs and exhibit complex spatio-temporal coupling across multiple channels. Convolutional Neural Networks (CNNs) have become the predominant approach for EEG signal analysis, owing to their effectiveness in local feature extraction and compatibility with grid-like sensor topologies. Nevertheless, the locality as-sumption inherent in conventional CNN s restricts their ability to capture functional connectivity and dynamic dependencies between spatially distant channels. To address this limitation, we propose NLCTCN, a novel non-local Temporal Convolutional Network that leverages a hierarchical greedy strategy to identify and exploit long-range correlations in multi-channel time series. We further introduce a new fusion scheme, integrated into an end-to-end lightweight CNN architecture to effectively combine these non-local interactions and optimize their configurations for improved predictive performance. Experimental results are presented on 10 real-world EEG datasets. These datasets cover human physiology, cognitive tasks, and clinical applications. The results show that NLCTCN significantly outperforms state-of-the-art methods. On average, NLCTCN achieves an accuracy improvement of 7.5 %. These results validate the effectiveness and superiority of the proposed approach in modeling non-local spatio-temporal dynamics under data-scarce and multi-channel conditions.
Han Zhang 0053, Lanqing Yang, Zechen Li 0005, Leping Yang, Yiheng Bian, Dian Ding, Leyu Jiang, Yi-Chao Chen 0001, Guangtao Xue
BIBM6
2025 M2SILENT: Enabling Multi-user Silent Speech Interactions via Multi-directional Speakers in Shared Spaces
abstract
We introduce M 2 Silent, which enables multi-user silent speech interactions in shared spaces using multi-directional speakers.Ensuring privacy during interactions with voice-controlled systems presents significant challenges, particularly in environments with multiple individuals, such as libraries, offices, or vehicles.M 2 Silent addresses this by allowing users to communicate silently, without producing audible speech, using acoustic sensing integrated into directional speakers.We leverage FMCW signals as audio carriers, simultaneously playing audio and sensing the user's silent speech.
Juntao Zhou, Dian Ding, Yijie Li 0002, Yu Lu 0022, Yida Wang 0007, Yongzhao Zhang, Yi-Chao Chen 0001, Guangtao Xue
CHI2
2025 AnchorAttention: Difference-Aware Sparse Attention with Stripe Granularity
abstract
Large Language Models (LLMs) with extended context lengths face significant computational challenges during the pre-filling phase, primarily due to the quadratic complexity of self-attention. Existing methods typically employ dynamic pattern matching and block-sparse low-level implementations. However, their reliance on local information for pattern identification fails to capture global contexts, and the coarse granularity of blocks leads to persistent internal sparsity, resulting in suboptimal accuracy and efficiency. To address these limitations, we propose AnchorAttention, a difference-aware, dynamic sparse attention mechanism that efficiently identifies critical attention regions at a finer stripe granularity while adapting to global contextual information, achieving superior speed and accuracy. AnchorAttention comprises three key components: (1) Pattern-based Anchor Computation, leveraging the commonalities present across all inputs to rapidly compute a set of near-maximum scores as anchor; (2) Difference-aware Stripe Sparsity Identification, performing difference-aware comparisons with anchor to quickly obtain discrete coordinates of significant regions in a stripe-like sparsity pattern; (3) Fine-grained Sparse Computation, replacing the traditional contiguous loading strategy with a discrete key-value loading approach to maximize sparsity rates while preserving hardware computational potential. Additionally, we integrate the identification strategy into a single operator to maximize parallelization potential. With its finer-grained sparsity strategy, AnchorAttention achieves higher sparsity rates at the same recall level, significantly reducing computation time. Compared to previous state-of-the-art methods, at a text length of 128k, it achieves a speedup of 1.44\times while maintaining higher recall rates.
Guoliang Zhu, Dian Ding, Yiming Zhang 0003
EMNLP5
2025 Hey Hey, My My, Skewness Is Here to Stay: Challenges and Opportunities in Cloud Block Store Traffic
abstract
Elastic Block Storage (EBS) has a pivotal role in modern data center infrastructure, providing reliable, high-performance and flexible block storage service to users. In Alibaba Cloud, EBS is the most widely used service and has been supporting the operation of millions of virtual disks. However, even with layers of load balancing and caching, we still observe significant traffic skewness across the EBS stack. This motivates us to comprehensively investigate symptoms and root causes behind the traffic patterns and, more importantly, explore the fixes for the identified issues.
Erci Xu, Yuandong Hong, Changsheng Niu, Lingjun Zhu, Jinnian He, Weidong Zhang 0011, Qiuping Wang, Changhong Wang 0005, Xinqi Chen, Guangtao Xue, Yi-Chao Chen 0001, Dian Ding
EuroSys16
2025 Monosulfide: A Sharded PoW Blockchain System with Secure Adaptive Mining Power Allocation
Guangtao Xue, Shengyun Liu, Jiahao Qi, Dian Ding
ICA3PP (6)6
2025 AMSER: Accelerate Mobile Speech Emotion Recognition with Signal Compression
abstract
Speech-based interaction systems are widely used in mobile devices like smartphones. With advances in deep neural networks, tasks such as speech emotion recognition (SER) enhance these systems’ user-friendliness. However, deploying SER models on mobile devices is challenging due to their complexity and computational demands. While pruning can reduce complexity, it often compromises accuracy, and hardware accelerators like FPGAs are difficult to integrate into mobile devices. This paper proposes AMSER, a real-time speech emotion recognition framework using signal compression and task offloading. AMSER utilizes logarithmic Mel-filter bank coefficients (Fbank) and singular value decomposition (SVD) for feature extraction and compression. The compressed signal is only 6.25% of the original size, achieving 2.24x faster transfer rates and 55.35% energy savings compared to raw audio transmission. Despite the compression, the features preserve key audio information for text and emotion recognition, performed server-side. Experiments show a WER of 4.68% (Librispeech), 10.69% (CommonVoice), and 69.83% emotion recognition accuracy (IEMOCAP).
Yu Lu 0022, Dian Ding, Han Zhang 0053, Lanqing Yang, Yi-Chao Chen 0001, Guangtao Xue
ICASSP3
2025 LLM4Load: An LLM Prompt-Driven Framework for Multi-Scale Workload Prediction
abstract
As application migration to the cloud becomes the mainstream way of application deployment, accurate workload prediction is critical to ensure the quality of service (QoS) and cost-efficiency of the applications and meet service level agreements (SLAs) with users. Short-term workload prediction can handle workload fluctuations over a short duration, while long-term prediction can capture trend and periodic changes for workload. However, existing studies are unable to deal with long-term forecasts in multiple scales effectively; moreover, new containers lack historical data, leading to inaccurate prediction, while existing methods perform poorly due to a lack of generalization ability. To tackle these problems, we propose LLM4Load, an LLM Prompt-Driven Framework for Multi-Scale Workload Prediction. Firstly, we design a structured prompt with dataset introduction, task description, and workload features characterization to extract multi-scale features. Secondly, we introduce a cross-modality alignment mechanism combined with label embedding to further enhance performance. Leveraging the generalization capability of LLM, we also solve the container cold-start problem. The abundant experiments on four real cloud workload datasets validate the effectiveness of LLM4Load. For multi-scale workload prediction, LLM4Load improves up to 42.40%. For container cold-start scenarios, LLM4Load improves up to 93.72%. These results highlight the potential of LLM4Load to drive dynamic and efficient resource management in modern cloud systems.
Zeyuan Ding, Dian Ding, Han Zhang 0053, Jiannong Cao 0001, Guangtao Xue
ICWS2
2025 Breaking the Mainchain Barrier of Blockchain Sharding Architecture for Federated Learning
abstract
Blockchain enhances the robustness and user engagement of Federated Learning (FL) systems but fails to meet the throughput and real-time requirements for model transmission. While sharding architectures improve system throughput, the latency introduced by mainchain model transmission remains a performance bottleneck, compromising the QoS of FL systems. In this paper, we propose a Mainchain-Free Sharding architecture, MFSChain, featuring an adaptive sharding mechanism based on hierarchical clustering. This mechanism improves shard model performance by eliminating the need for mainchain aggregation (i.e., shard-level global models). We also introduce the Federated Learning State Tree (FLS-Tree) for client management and state migration without a mainchain, alongside a lightweight storage scheme, LiFLS-Tree. Through theoretical analysis and extensive simulations, we demonstrate that MFSChain outperforms traditional blockchain and sharding architectures. Specifically, MFSChain reduces client waiting time by 17% and 24%, increases average model accuracy by 2.5% to 10% compared to traditional global models, and boosts throughput by$464 \times$while reducing transaction processing latency by 99%.
Jiahao Qi, Dian Ding, Han Zhang 0053, Yi-Chao Chen 0001, Jiong Lou, Jiadi Yu, Qiaoling Xiao, Jie Li 0002, Jiannong Cao 0001, Guangtao Xue
IWQoS2
2025 Bridge: Enabling BLE Direction Finding Feature Compatible with All Bluetooth Devices
abstract
Bluetooth-based location services have experienced significant growth over the past decades. RSSI-based techniques using beacons only provide meters-level accuracy. Angular-based approaches rely on customized antenna arrays, introducing high costs and limited usability. In 2020, Bluetooth Special Interest Group (Bluetooth SIG) released version 5.1, integrating Angle of Arrival (AoA) estimation to enable direction finding capabilities, which has the potential to improve localization across various fields, including logistics and industry. However, more than 4.1 billion devices (68% of the total) still do not support the direction finding feature. To address this issue and ensure backward compatibility, we proposed Bridge, a solution that leverages an additional trigger node (referred to as Trigger) to make the direction finding feature compatible with all Bluetooth devices without requiring modifications to existing hardware or firmware. The Trigger mimics communication behaviors with both locators and targets simultaneously by sending a nesting packet. Subsequently, processes and algorithms are delicately designed to estimate AoA. Bridge also supports large-scale deployment through dynamic packet flow switching, enabling it to handle concurrent targets and manage handover with a consistent operation pattern. We implemented and evaluated Bridge in real-world scenarios. The system achieved an average localization error of 33.4cm while extending the direction-finding feature to 10 target devices of different Bluetooth versions, indicating the effectiveness of Bridge.
Runting Zhang, Yijie Li 0002, Dian Ding, Yi-Chao Chen 0001, Yida Wang 0007, Dongyao Chen, Jiadi Yu, Guangtao Xue
MobiCom3
2025 Poster: Enabling BLE Direction Finding Feature Compatible with All Bluetooth Devices
abstract
BLE direction finding provides high-accuracy localization based on Angle-of-Arrival (AoA), but this feature is only available on BLE 5.1+ devices. Billions of existing Bluetooth devices are excluded from direction finding indoor localization systems. We present Bridge that enables direction finding for all Bluetooth versions without any hardware or firmware modifications. Bridge introduces a novel Trigger that mimics communication behaviors of both locators and targets, allowing the locator to extract AoA information from originally unsupported devices. We implement Bridge on COTS direction finding system and evaluate it on 10+ BLE devices, achieving a median localization error of 33.4cm.
Runting Zhang, Yijie Li 0002, Dian Ding, Yi-Chao Chen 0001
MobiCom3
2025 SADIF: Spoofing Attack on BLE Direction Finding Based Localization System
abstract
Bluetooth Low Energy (BLE) direction finding, a feature introduced in BLE version 5.1, enables precise localization through Angle of Arrival (AoA) estimation. However, this advancement introduces new risk to BLE direction finding based localization system. Specifically, the AoA estimation based on phase sampling of constant-tone-extension (CTE) is susceptible to the signal injection attack. This paper presents SaDiF, a feasible spoofing attack mechanism to mislead the locators into mistaking the positioning result as a continuous path. By eavesdropping on BLE packets and injecting attack signals containing pre-designed disturbing phase shift, SaDiF subtly alters the AoA estimation without detection, thus interfere the localization results. Moreover, SaDiF address the challenges posed by hardware imperfections by proposing an injection timing optimization to improve attack robustness. Extensive experiments demonstrates the effectiveness of SaDiF in successfully attacking multiple BLE targets in real-time scenarios. In conclusion, our findings reveal critical security risks in BLE direction finding feature and provide insights into strengthening its defenses.
Runting Zhang, Yijie Li 0002, Dian Ding, Hao Pan 0003, Yongzhao Zhang, Yi-Chao Chen 0001, Xiaoyu Ji 0001, Jiadi Yu, Guangtao Xue
MobiHoc3
2025 High-resolution mmWave Imaging using Metasurface and Diffusion
Yida Wang 0007, Yu Lu 0022, Yifei Shen 0004, Lili Qiu, Zeyuan Lai, Yi-Chao Chen 0001, Hao Pan 0003, Juntao Zhou, Dian Ding, Guangtao Xue, Qian Zhang 0001
MobiSys10
2025 MODepth: Benchmarking Mobile Multi-frame Monocular Depth Estimation with Optical Image Stabilization
abstract
This paper presents MODepth, a multi-frame monocular depth estimation system based on the controlled motion of an optical image stabilization (OIS) module. By actively injecting acoustic signals, we induce regular translational movements of the OIS lens, resulting in controllable camera pose changes and simplifying inter-frame pose estimation. Leveraging multi-frame images captured under OIS-controlled lens movements, we design a high-precision depth estimation network, MODNet, and introduce the principal point offset estimation module and pose estimation modules to fully exploit geometric information across frames. To validate the effectiveness of our approach, we collect a new dataset MODdata with 1100 samples in nearly 220 indoor scenarios and benchmark our model as an OIS-based multi-frame depth estimation method, comparing it to ground truth obtained from a depth sensor and other state-of-the-art monocular depth estimation algorithms. Our method achieves competitive or superior performance compared to fully supervised baselines, reaching an RMSE of 0.439, which outperforms all evaluated methods, demonstrating that self-supervised fine-tuning with OIS-induced parallax is a viable alternative to ground-truth supervision. Code and dataset are available at: https://github.com/liangjindeamo-yuer/MODEPTH
Yu Lu 0022, Hao Pan 0003, Dian Ding, Jiatong Ding, Yongjian Fu 0004, Yi-Chao Chen 0001, Ju Ren 0001, Guangtao Xue
SIGGRAPH Asia3
2025 Amser+: Accelerating Mobile Speech Emotion Recognition in IoT Environments With Mel Feature Compression
abstract
Speech-based interaction systems are widely used in mobile devices like smartphones. With advances in deep neural networks, tasks such as speech emotion recognition (SER) enhance these systems user-friendliness. However, deploying SER models on mobile devices is challenging due to their complexity and computational demands. While pruning can reduce complexity, it often compromises accuracy, and hardware accelerators like FPGAs are difficult to integrate into mobile devices. This paper proposes Amser+, a real-time speech emotion recognition framework using signal compression and task offloading. Amser+utilizes logarithmic Mel-filter bank coefficients (Fbank) and singular value decomposition (SVD) for feature extraction and compression. The compressed signal is only 6.25% of the original size, achieving 2.24× faster transfer rates and 55.35% energy savings compared to raw audio transmission. Despite the compression, the features preserve key audio information for text and emotion recognition, performed server-side. Experiments show a WER of 4.68% (Librispeech), 10.69% (CommonVoice), and 72.85% emotion recognition accuracy (IEMOCAP).
Yu Lu 0022, Dian Ding, Yijie Li 0002, Yongzhao Zhang, Lanqing Yang, Yi-Chao Chen 0001, Guangtao Xue
IEEE Internet Things J.3
2025 MasterPlan: A Reinforcement Learning Based Scheduler for Archive Storage
abstract
With the sheer volume of data in today’s world, archive storage systems play a significant role in persisting the cold data. Due to stringent cost concerns, one popular design is to organize disks into groups and periodically switch them to be powered on for serving user requests. Scheduling thus becomes critical for both CapEx and performance. Unfortunately, field results indicate that existing schedulers can be often suboptimal. Our further analysis suggests that the main reason is the mismatch between the ever-changing workloads and the fixed set of coarsely-configured parameters in current heuristic-based schedulers. In this article, we propose MasterPlan , a reinforcement learning (RL) based scheduler for archive storage systems. By identifying the unique characteristics of archive storage service, we design a state space and reward function for the RL agent. MasterPlan includes a continuous action encoding approach to guarantee efficient exploration, and a meta adaptation module to extract features of workload series. Experiments show that MasterPlan can achieve 1.25× throughput, 2.16× 99 th latency and 1.47× power draw improvement compared to existing solutions.
Xinqi Chen, Erci Xu, Dengyao Mo, Ruiming Lu, Dian Ding, Guangtao Xue
ACM Trans. Archit. Code Optim.6
2025 TouchHBC: Touch-Based Human Body Communication via Leakage Current
abstract
Wearable devices, including smartwatches, are increasingly popular among consumers due to their user-friendly services. However, transmitting sensitive data like social media messages and payment QR codes via commonly used low-power Bluetooth exposes users to privacy breaches and financial losses. This study introducesTouchHBC, a secure and reliable communication scheme leveraging a smartwatch's built-in electrodes. This system establishes a touch-based human communication system utilizing a laptop's leakage current. As the transmitting device, the laptop modulates this current via the CPU. Simultaneously, the smartwatch, equipped with built-in electrodes, captures the current traversing the human body and decodes it. The modulation and decoding processes involve techniques such as amplitude modulation, variational mode decomposition, channel estimation, and retransmission mechanisms.TouchHBCfacilitates communication between laptops and smartwatches. Real-world tests demonstrate that our prototype achieves a throughput of$19.83bps$. Moreover,TouchHBCoffers the potential for enhanced interaction, including improved gaming experiences through vibration feedback and secure touch login for smartwatch applications by synchronizing with a laptop. Furthermore, the system can be integrated with high-throughput communication protocols such as Bluetooth, enhancing its scalability while maintaining a strong foundation of security.
Dian Ding, Hao Pan 0003, Yongzhao Zhang, Yijie Li 0002, Yu Lu 0022, Yi-Chao Chen 0001, Guangtao Xue
IEEE Trans. Mob. Comput.1
2025 SwiftTrack+: Fine-Grained and Robust Fast Hand Motion Tracking Using Acoustic Signal
abstract
Acoustic tracking technology, leveraging the ubiquitous presence of speakers and microphones in commercial off-the-shelf (COTS) mobile devices, has become a versatile tool across various applications. However, current phase-based acoustic tracking methods encounter significant limitations in tracking fast movements, thereby restricting their practical utility. This paper identifies three practical challenges to enable fast hand motion tracking using acoustic signals: 1) high mobility, 2) low signal-to-noise ratio (SNR), and 3) variations in hardware frequency response. The high mobility introduces Doppler shift and phase ambiguity which is the primary cause of failure in fast movement tracking, while the latter two factors can further impair the tracking performance in practical scenarios involving high mobility. To address the high mobility issue, we effectively compensate the Doppler shift in the Channel Impulse Response (CIR) for better selection of channel taps and then propose a novel phase derivative approach to mitigate the phase ambiguity. To enhance the real-world robustness, we integrate multiple algorithms including an SNR enhancement algorithm inspired by time-domain beamforming and a hardware frequency response compensation approach that addresses both amplitude and phase distortions. Additionally, an LSTM-based distance reconstruction algorithm is further implemented to correct residual phase noise. Implemented on Android platforms under the name SwiftTrack+, our system demonstrates superior performance in tracking fast movements. Through extensive evaluations, SwiftTrack+ proves its efficacy across diverse scenarios, significantly broadening the scope and reliability of acoustic tracking applications.
Yongzhao Zhang, Hao Pan 0003, Dian Ding, Yi-Chao Chen 0001, Lili Qiu, Guangtao Xue, Ting Chen 0002, Xiaosong Zhang 0001
IEEE Trans. Netw.3
2024 CarbonNet: Enterprise-Level Carbon Emission Prediction with Large-Scale Datasets
Jinghua Tang, Lanqing Yang, Yuqiao Pei, Dian Ding, Yu Lu 0022, Guangtao Xue
ICIC (12)6
2024 DASIV: Directional Acoustic Sensing based Intelligent Vehicle Interaction System
abstract
With the increase in motor vehicles, more convenient and accurate interactions are expected while retaining a high standard of safe driving. However, complex and dynamic vehicle environments challenge sensing tasks such as breathing monitor and hand gesture recognition. In this paper, we propose DASIV, which utilizes the highly directional nature of ultrasonic signals to achieve fine-grained directional acoustic sensing in vehicle environments. Due to air nonlinearity, the system enables synchronized directional acoustic communication to transmit information (e.g., navigation) to the driver without affecting other passengers. By optimizing the frequency of the Frequency Modulated Continuous Wave (FMCW) signals, DASIV avoids mutual interference between the sensing and communication signals and achieves breathing detection and hand gesture recognition for the driver. Specifically, the system extracts breathing-induced weak thoracic bullying through the signal phase, captures and analyses breathing patterns using bandpass and Gaussian filters, and develops a breathing model. Then, the system defines 10 interaction hand gestures to meet daily interaction needs, uses spectral features to mine complex and fast hand movement features, and proposes a hand gesture recognition model. Extensive experiments in real environments show that DASIV achieves high-precision breathing monitor (Pearson correlation coefficient of 0.89) and hand gesture recognition (Precision of 91.7%).
Dinghua Zhao, Juntao Zhou, Dian Ding, Yu Lu 0022, Yijie Li 0002, Yi-Chao Chen 0001, Guangtao Xue
IPCCC3
2024 MuDiS: An Audio-independent, Wide-angle, and Leak-free Multi-directional Speaker
abstract
This paper introduces a novel multi-directional speaker, named MuDiS, which utilizes a parametric array to generate highly focused sound beams in multiple directions. The system capitalizes on air nonlinearity to reproduce sound from ultrasounds, successfully overcoming challenges inherent in traditional parametric arrays, such as transducer size and wavefront shape. It supports three important features simultaneously: independent beams, wide-angle digital steering, and unintended leakage suppression. To address these challenges, we designed a specialized cell structure that connects ultrasonic transducers, redirecting an approximately omnidirectional wavefront with optimal interspacing. An optimization-based algorithm is developed to minimize unintended leakages, and a nonlinear distortion reduction scheme is proposed to enhance sound quality. The paper showcases a prototype demonstrating the system's capabilities as a multidirectional speaker with a wide sound projection angle. Experimental results validate the effectiveness of our approach. The proposed multi-beam projection system rivals the performance of commercially available single-beam projection directional speakers, and improved steering angle and sound fidelity compared to multi-beamforming performance using traditional parametric arrays.
Yijie Li 0002, Juntao Zhou, Dian Ding, Yi-Chao Chen 0001, Lili Qiu, Jiadi Yu, Guangtao Xue
MobiCom3
2024 M3Cam: Extreme Super-resolution via Multi-Modal Optical Flow for Mobile Cameras
abstract
The demand for ultra-high-resolution imaging in mobile phone photography is continuously increasing. However, the image resolution of mobile devices is typically constrained by the size of the CMOS sensor. Although deep learning-based super-resolution (SR) techniques have the potential to overcome this limitation, existing SR neural network models require large computational resources, making them unsuitable for real-time SR imaging on current mobile devices. Additionally, cloud-based SR systems pose privacy leakage risks. In this paper, we propose M3Cam, an innovative and lightweight SR imaging system for mobile phones. M3Cam can ensure high-quality 16× SR image (4× in both height and width) visualization with almost negligible latency. In detail, we utilize an optical image stabilization (OIS) module for lens control and introduce a new modality of data, namely gyroscope readings, to achieve high-precision and compact optical flow estimation modules. Building upon this concept, we design a multi-frame-based SR model utilizing the Swin Transformer. Our proposed system can generate a 16× SR image from four captured low-resolution images in real-time, with low computational load, low inference latency, and minimal reliance on runtime RAM. Through extensive experiments, we demonstrate that our proposed multi-modal optical flow model significantly enhances pixel alignment accuracy between multiple frames and delivers outstanding 16× SR imaging results under various shooting scenarios. Code and dataset are available at: https://github.com/liangjindeamo-yuer/M3CAM
Yu Lu 0022, Dian Ding, Hao Pan 0003, Yongjian Fu 0004, Feitong Tan, Yi-Chao Chen 0001, Guangtao Xue, Ju Ren 0001
SenSys2
2024 HandPad: Make Your Hand an On-the-go Writing Pad via Human Capacitance
abstract
The convenient text input system is a pain point for devices such as AR glasses, and it is difficult for existing solutions to balance portability and efficiency. This paper introduces HandPad, the system that turns the hand into an on-the-go touchscreen, which realizes interaction on the hand via human capacitance. HandPad achieves keystroke and handwriting inputs for letters, numbers, and Chinese characters, reducing the dependency on capacitive or pressure sensor arrays. Specifically, the system verifies the feasibility of touch point localization on the hand using the human capacitance model and proposes a handwriting recognition system based on Bi-LSTM and ResNet. The transfer learning-based system only needs a small amount of training data to build a handwriting recognition model for the target user. Experiments in real environments verify the feasibility of HandPad for keystroke (accuracy of 100%) and handwriting recognition for letters (accuracy of 99.1%), numbers (accuracy of 97.6%) and Chinese characters (accuracy of 97.9%).
Yu Lu 0022, Dian Ding, Hao Pan 0003, Yijie Li 0002, Juntao Zhou, Yongjian Fu 0004, Yongzhao Zhang, Yi-Chao Chen 0001, Guangtao Xue
UIST2
2023 LeakThief: Stealing the Behavior Information of Laptop via Leakage Current
abstract
Smart devices are proliferating in every aspect of our lives, providing convenience but also exposing us to the risk of information leakage at any moment. Attackers can monitor the user and infer private information such as the personality and preferences by stealing the behavior information. In this paper, we investigated the potential threat of information stealing via the leakage current of laptop and electrodes in wearable devices (e.g. smart watches and bracelets). Specifically, the leakage current in the laptop adapter can flow from the metal casing into the human body and be collected by electrodes in wearable devices when the user is using a laptop with a metal casing (e.g. MacBook). We verified the correlation between leakage current and working states of the laptop, where different operations corresponding to different CPU instructions can generate different leakage currents. Based on this, we propose LeakThief, the system consists of three components, leakage current detection, application operation detection and application recognition. The experiments in real-world environment demonstrated that the proposed system is able to recognize 10 common applications with high accuracy, including launching-based (97.5%) and in-application operation-based recognition (83.8%).
Dian Ding, Yi-Chao Chen 0001, Xiaoyu Ji 0001, Guangtao Xue
SECON1
2023 Handwriting Recognition System Leveraging Vibration Signal on Smartphones
abstract
The efficiency of human-computer interaction is greatly hindered by the small size of the touch screens on mobile devices, such as smart phones and watches. This has prompted widespread interest in handwriting recognition systems, which can be divided into active and passive systems. Active systems require additional hardware devices to perceive movements of handwriting or the tracking accuracy is not adequate for handwriting recognition. Passive methods use the acoustic signal of pen rubbing and are susceptible to environmental noise (above 60$dB$). This paper presents a novel handwriting recognition system based on vibration signals detected by the built-in accelerometer of smartphones. The proposed scheme is implemented in three stages: signal segmentation, signal recognition, and word suggestion.VibWriteris highly resistant to interferences since the normal environmental noise (below 70$dB$) will not cause the vibration of the accelerometer. Extensive experiments demonstrated the efficacy of the system in terms of accuracy in letter recognition (75.3%), word recognition (86.4%) and number recognition (79%) in a variety of writing positions under a variety of environmental conditions.
Dian Ding, Lanqing Yang, Yi-Chao Chen 0001, Guangtao Xue
IEEE Trans. Mob. Comput.1
2021 VibWriter: Handwriting Recognition System based on Vibration Signal
abstract
The efficiency of human-computer interaction is greatly hindered by the small size of the touchscreens on mobile devices, such as smart phones and watches. This has prompted widespread interest in handwriting recognition systems, which can be divided into active and passive systems. Active systems require additional hardware devices to perceive movements of handwriting or the tracking accuracy is not adequate for hand-writing recognition. Passive methods use the acoustic signal of pen rubbing and are susceptible to environmental noise (above 60dB). This paper presents a novel handwriting recognition system based on vibration signals detected by the built-in accelerometer of smart phones. VibWriter is highly resistant to interference since the normal environmental noise will not cause the vibration of the accelerometer. Extensive experiments demonstrated the efficacy of the system in terms of accuracy in letter recognition (76.15%) and word recognition (88.14%) when dealing with words of various lengths written by various users in a variety of writing positions under a variety of environmental conditions.
Dian Ding, Lanqing Yang, Yi-Chao Chen 0001, Guangtao Xue
SECON1
2020 MagPrint: Deep Learning Based User Fingerprinting Using Electromagnetic Signals
abstract
Understanding the nature of user-device interactions (e.g., who is using the device and what he/she is doing with it) is critical to many applications including time management, user profiles, and privacy protection. However, in scenarios where mobile devices are shared among family members or multiple employees in a company, conventional account-based statistics are not meaningful. This poses an even bigger problem when dealing with sensitive data. Moreover, fingerprint readers and front-facing cameras were not designed to continuously identify users. In this study, we developed MagPrint, a novel approach to fingerprint users based on unique patterns in the electromagnetic (EM) signals associated with the specific use patterns of users. Initial experiments showed that time-varying EM patterns are unique to individual users. They are also temporally and spatially consistent, which makes them suitable for fingerprinting. MagPrint has a number of advantages over existing schemes: i) Non-intrusive fingerprinting, ii) implementation using a small and easy-to-deploy device, and iii) high accuracy thanks to the proposed classification algorithm. In experiments involving 30 users, MagPrint achieves 94.3% accuracy in classifying users from these traces, which represents an 10.9% improvement over the state-of-the-art classification method.
Lanqing Yang, Yi-Chao Chen 0001, Hao Pan 0003, Dian Ding, Guangtao Xue, Linghe Kong, Jiadi Yu, Minglu Li 0001
INFOCOM4