VLDB 2026 Research / reviewers in the wild / expert
Di Duan
dblp:345/2464
· DBLP profile ↗
15ranked-venue papers
4as first author
15since 2021 · last 2026
0000-0003-4184-0762ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 11 · 2 first-author · 11 since 2021Human-computer interaction and ubiquitous computing · 3 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | COACH: Adaptive Robust Human-Robot Collaboration for Efficient Smart ManufacturingabstractModern smart manufacturing pipelines have pervasively collaborated human workers, mobile robots, and industrial Internet of Things (IIoT) in shared workspaces for versatile production tasks. Despite the promising capacity of individual entities, the performance of these IIoT systems largely relies on pipeline coordination, i.e., task dispatching between humans and robots, which is particularly challenging under heterogeneous physical constraints and complex environmental uncertainties. Nonetheless, existing works either rely on traditional operation frameworks that lack scalability for large-scale complex production, or propose customized solutions for fixed agent models, overlooking the evolving nature of IIoT environments. To address these limitations, this paper proposes COACH, a human-robot collaborative manufacturing system that enables robust constraint-aware coordination across humans, robots, and IIoT. Specifically, COACH designs a scalable contextual encoder to represent the evolving relationships among human and robot agents in dynamic heterogeneous graphs. With that, a novel experience-driven task dispatcher is developed, enabling both high-performance and computation-efficient policy generation concerning the status of IIoT. To accommodate changing human fatigue and pipeline scales, COACH further develops a curriculum-enhanced reinforcement learning module for efficient dispatcher adaptation. Extensive evaluations using both synthetic testbeds and real-world manufacturing datasets demonstrate that COACH improves the feasible ratio of manufacturing pipelines by up to 27.4% and achieves up to 13.9% improvement in time efficiency compared to competing baselines across diverse job scales and environmental settings. Hui Wang 0011, Liekang Zeng, Zhiwen Yu 0001, Yao Zhang 0005, Di Duan, Mu Yuan, Bin Guo 0001, Guoliang Xing |
SenSys | 5 |
| 2026 | SwiftChannel: Algorithm-Hardware Co-Design for Deep Learning-Based 5G Channel EstimationabstractChannel estimation is crucial in 5G communication networks for optimizing transmission parameters and ensuring reliable, high-speed communication. However, the use of multiple-input and multiple-output (MIMO) and millimeter-wave (mmWave) in 5G networks presents challenges in achieving accurate estimation under strict latency requirements on resource-limited hardware platforms. To address these challenges, we proposeSwiftChannel, an algorithm-hardware co-design framework that integrates a hardware-friendly deep learning-based channel estimator with a dedicated accelerator. Our approach employs a convolutional neural network enhanced with a parameter-free attention mechanism, which effectively reconstructs full-resolution spatial-frequency domain channel matrices from low-resolution least squares (LS) estimates. We further develop a multi-stage model compression pipeline combining knowledge distillation, convolution re-parameterization, and quantization-aware training, resulting in substantial model size reduction with negligible accuracy loss. The hardware accelerator, implementing the compressed model and the LS estimator on FPGA platforms using High-level Synthesis (HLS), features a fine-grained pipeline architecture and optimized dataflow strategies. Tested on a Zynq UltraScale+ RFSoC, the accelerator achieves sub-millisecond latency, providing up to 24x speed-up and over 33x improvement in energy efficiency compared to GPU-based solutions. Extensive evaluations demonstrate that the proposed design generalizes not only across various noise levels and user mobilities, but also to a variety of unseen channel profiles, outperforming state-of-the-art baselines. By unifying algorithmic innovation with hardware-aware design, our work presents a future-proof channel estimation solution for 5G MIMO systems. The source codes for the dataset synthesis, deep learning algorithm, and HLS-based FPGA design are accessible via GitHub. Shengzhe Lyu, Yuhan She, Di Duan, Tao Ni 0003, Yu Hin Chan, Chengwen Luo 0001, Ray C. C. Cheung, Weitao Xu |
IEEE Trans. Mob. Comput. | 3 |
| 2025 | iRadar: Synthesizing Millimeter-Waves from Wearable Inertial Inputs for Human Gesture Sensing
Huanqi Yang, Mingda Han, Di Duan, Tianxing Li 0001, Weitao Xu |
INFOCOM | 4 |
| 2025 | Grape: Efficient Spatiotemporal Prediction Services with Stale Sensing StreamsabstractEmerging cyber-physical systems have embraced a large number of IoT devices spanning geo-distributed, which generate and consume massive volumes of data continuously. Accurate and timely spatiotemporal predictions (STP) over these streaming sensor data are critical and, in growing demand, ubiquitous across various edge scenarios such as traffic flow forecasting. Towards that, recent advanced systems have developed sophisticated optimizations among STP pipelines, aiming at optimal prediction performance. However, based on our empirical studies in real-world settings, we identify a previously overlooked bottleneck of end-to-end STP performance: data staleness. To mitigate this issue, in this work, we investigate a new task, namely stream interception, which deliberately terminates the acceptance of incoming sensor data and anticipates model execution with imputed missing features. We propose a novel dynamic interception strategy to determine the time slot to exit waiting and present Grape, an STP system that implements it with practical system designs. Extensive evaluations on real-world traces show that Grape can strike a superior tradeoff between prediction accuracy and serving latency, achieving 1.69-1.90× speedup against traditional all-waiting baselines across various STP services with high prediction accuracy on par with offline optimal cases. Liekang Zeng, Shengyuan Ye, Mu Yuan, Di Duan, Xu Chen 0004, Guoliang Xing |
RTSS | 5 |
| 2025 | Argus: Multi-View Egocentric Human Mesh Reconstruction Based on Stripped-Down Wearable mmWave Add-onabstractIn this paper, we propose Argus, a wearable add-on system based on stripped-down (i.e., compact, lightweight, low-power, limited-capability) mmWave radars. It is the first to achieve egocentric human mesh reconstruction in a multi-view manner. Compared with conventional frontal-view mmWave sensing solutions, it addresses several pain points, such as restricted sensing range, occlusion, and the multipath effect caused by surroundings. To overcome the limited capabilities of the stripped-down mmWave radars (with only one transmit antenna and three receive antennas), we tackle three main challenges and propose a holistic solution, including tailored hardware design, sophisticated signal processing, and a deep neural network optimized for high-dimensional complex point clouds. Extensive evaluation shows that Argus achieves performance comparable to traditional solutions based on high-capability mmWave radars, with an average vertex error of 6.5 cm, solely using stripped-down radars deployed in a multi-view configuration. It presents robustness and practicality across conditions, such as with unseen users and different host devices. Di Duan, Shengzhe Lyu, Mu Yuan, Hongfei Xue, Tianxing Li 0001, Weitao Xu, Kaishun Wu, Guoliang Xing |
SenSys | 1 |
| 2025 | SCX: Stateless KV-Cache Encoding for Cloud-Scale Confidential Transformer ServingabstractTransformer models have revolutionized fields like natural language processing and computer vision but face privacy concerns in sensitive applications such as medical diagnostics. Existing confidential serving methods, including cryptography-based, memory isolation-based, and access control-based, offer trade-offs between privacy and efficiency but often struggle with high latency or hardware dependencies. This work proposes stateless KV-cache encoding (SCX), a novel framework that encodes the intermediate key-value cache during Transformer inference using user-controlled keys. SCX ensures that the cloud can neither recover the input nor independently complete the next token prediction, effectively preserving privacy. By introducing efficient encoding and decoding schemes, SCX addresses communication complexity and attack vulnerabilities while ensuring zero loss of inference quality. Experiments on large Transformer models demonstrate that SCX achieves lower latency (e.g., 36ms for LLaMA-7B), outperforming state-of-the-art cryptography and memory isolation methods by orders of magnitude. Moreover, SCX can complementarily work with advanced KV-cache management techniques to further enhance KV-cache communication efficiency by 85%, marking a significant step toward practical, privacy-preserving large Transformer serving. Mu Yuan, Lan Zhang 0002, Liekang Zeng, Siyang Jiang, Bufang Yang, Di Duan, Guoliang Xing |
SIGCOMM | 6 |
| 2025 | PricoEye: The Eye of Primary Colors for Fast and Convenient 3D Reconstruction of Fine-grained Palmprint on Smartphones
Di Duan, Kaicheng Xiao, Lixing He, Wei Gao 0006, Guoliang Xing |
UIST | 1 |
| 2025 | Mitigating Tail Latency for On-Device Inference With Load-Balanced Heterogeneous ModelsabstractServing machine learning models on edge, mobile, and embedded devices places stringent requirements on inference latency. From operating a real enterprise service, we observed that even a fully optimized model could lead to severe violations of latency objectives when the load surges. A straightforward and mature approach is to auto-scale multiple models to balance the load. However, unlike cloud clusters, edge or mobile devices usually cannot afford to deploy multiple model replicas. Therefore, in this paper, we explore a new idea: in addition to the original model, we deploy one (or more) heterogeneous model(s) with much smaller resource overhead on the device, and perform load balancing among all models. We overcame the technical challenges posed by performance dynamics and developed InferRouter based on queuing theory. We implement and evaluate InferRouter on three real on-device inference systems, covering mobile sensing, video analytics, and natural language processing applications. Experimental results show that compared with strong baselines, InferRouter can decrease 85.2% P99 latency (5.8x faster) and improve 5.9% accuracy on the mobile workload. For a traffic video analytics task, InferRouter achieves 55.1% higher accuracy with zero deadline misses. InferRouter also shows its advantages in saving resources compared with auto-scaling and offloading approaches. Mu Yuan, Lan Zhang 0002, Di Duan, Liekang Zeng, Miaohui Song, Zichong Li, Guoliang Xing, Xiang-Yang Li 0001 |
IEEE Trans. Mob. Comput. | 3 |
| 2024 | RF-Egg: An RF Solution for Fine-Grained Multi-Target and Multi-Task Egg Incubation SensingabstractEggs and chickens serve as crucial animal-source proteins in our diets, making large-scale breeding egg incubation an essential undertaking. However, current solutions, i.e., vision-based and sensor-based methods, are primarily designed for egg fertility detection tasks under single-egg settings, which have not yet satisfied the goal of multi-target and multi-task sensing. In this paper, we propose RF-Egg, the first RF-based fine-grained multi-target and multi-task egg incubation sensing system with respect to sensing fertility, incubation status, and early mortality of chicken embryos. RF-Egg leverages the weak coupling effects of RFID tags when interacting with eggs, which induces different impedance changes of RFID tags with the incubation levels of eggs, thereby resulting in a variation of low-level phase readings of the backscatter signals. Regarding the challenge of multi-target profiling interference, we propose a multipath combating algorithm to extract the target-induced signal component based on the built signal model, and address non-uniformity issues across multiple tags. Moreover, we devise three unique feature maps tailored to each task, and then design an Multi-Task Triplet (MTT) network for multitasking. Our evaluation results based on 189 eggs show that RF-Egg achieves an accuracy of 94.4%, 96.1%, and 90.1% for the aforementioned three tasks when supporting 16 targets. Additionally, our extensive field study in a local egg hatchery suggests that RF-Egg presents the potential to be widely deployed in the modern poultry industry. Zehua Sun, Tao Ni 0003, Di Duan, Kai Liu 0008, Weitao Xu |
MobiCom | 4 |
| 2024 | F2Key: Dynamically Converting Your Face into a Private Key Based on COTS Headphones for Reliable Voice InteractionabstractIn this paper, we proposed F2Key, the first earable physical security system based on commercial off-the-shelf headphones. F2Key enables impactful applications, such as enhancing voiceprint-based authentication systems, reliable voice assistants, audio deepfake defense, and the legal validity of artifacts. The key idea of F2Key is to establish a stable acoustic sensing field across the user's face and embed the user's facial structures and articulatory habits into a user-specific generative model that serves as a private key. The private key can decrypt the Channel Impulse Response (CIR) profiles provided by the acoustic sensing field into an inferred spectrogram that can match the real one calculated from the corresponding speech, provided that the user's CIR-spectrogram mapping relationship is consistent with the one embedded in the generative model. Extensive experiments demonstrate that F2Key resists 99.9%, 96.4%, and 95.3% of speech replay attacks, mimicry attacks, and hybrid attacks, respectively. We discussed and evaluated F2Key from different perspectives, such as the health consideration and identical twins study, to show the practicality and reliability. Di Duan, Zehua Sun, Tao Ni 0003, Shuaicheng Li 0001, Xiaohua Jia, Weitao Xu, Tianxing Li 0001 |
MobiSys | 1 |
| 2024 | Demo: Myotrainer: Muscle-Aware Motion Analysis and Feedback System for In-Home Resistance TrainingabstractResistance training is widely incorporated in exercise programs, including in-home fitness and rehabilitation. However, improper motion patterns and muscle stimulation can undermine the safety of the subjects, making precise monitoring essential. Existing solutions primarily focus on correcting motion patterns with difficulties assessing muscle contraction levels. In this work, we introduce MyoTrainer, which provides muscle-aware motion descriptions and personalized feedback in natural language. Taking a person's exercise video as input, MyoTrainer first utilizes pose estimation models to capture motion sequences in real-time. A GCN-Former model has been developed for fine-grained motion analysis, which includes action recognition, incorrect movement pattern detection, and muscle contraction intensity estimation. Additionally, MyoTrainer integrates fitness and physiotherapeutic domain knowledge to deliver personalized, professional feedback. Extensive evaluations show that our system outperforms existing solutions in all recognition tasks and a survey indicates 88.9% of users find the generated feedback to be beneficial. Yuting He 0006, Xinyan Wang 0003, Mu Yuan, Di Duan, Doris Sau-Fung Yu, Guoliang Xing, Hongkai Chen 0001 |
SenSys | 4 |
| 2024 | Medusa3D: The Watchful Eye Freezing Illegitimate Users in Virtual Reality InteractionsabstractThe remarkable growth of Virtual Reality (VR) in recent years has extended its applications beyond entertainment to sectors including education, e-commerce, and remote communication. Since VR devices contain user's private information, user authentication becomes increasingly important. Current authentication systems in VR, such as password-based or static biometric-based methods, are either cumbersome to use or vulnerable to attacks such as shoulder surfing. To address these limitations, we propose Medusa3D, a challenge-response authentication system for VR based on reflexive eye responses. Unlike existing methods, reflexive eye responses are involuntary and effortless, offering a secure and user-friendly credential for authentication. We implement Medusa3D on an off-the-shelf VR and conduct evaluations with 25 participants. The evaluation results show that Medusa3D achieves 0.21% FAR and 0.13% FRR, demonstrating high security under various ocular conditions and resilience against attacks such as zero-effort attack, replay attack, and mimicry attack. A user study indicates that Medusa3D is user-friendly and well-adopted among participants. Aochen Jiao, Di Duan, Weitao Xu |
Proc. ACM Hum. Comput. Interact. | 2 |
| 2024 | Scenario-Adaptive Key Establishment Scheme for LoRa-Enabled IoV CommunicationsabstractIn recent years, the Internet of Vehicles (IoV) has experienced significant growth, but the lack of effective secret key establishment remains a security concern due to the dynamic and ad-hoc nature of IoV communications. Physical layer key generation has emerged as a promising solution for establishing a pair of cryptographic keys in a lightweight and information-theoretic secure manner. However, previous works have primarily focused on legacy communication technologies, such as Wi-Fi, ZigBee, and 5 G, which are limited to short-range IoV communications. With the emergence of Long-range (LoRa) communication technology, which features long-range, low power, and extremely low data rates, new challenges arise for key generation in long-range IoV scenarios. This paper presentsVehicle-Key, a secret key generation system designed to secure LoRa-enabled IoV communications.Vehicle-Keypresents an innovative scenario adaptive deep learning model that performs channel prediction and quantization concurrently while reducing the training cost through a data augmentation pipeline and enhancing the model's generalization using a domain-adaption method. Additionally, we propose a bloom filter-assisted autoencoder-based reconciliation method to significantly improve the key agreement rate. Comprehensive real-world experiments show thatVehicle-Keysurpasses the State-of-the-Art, achieving a 15.26%–50.35% improvement in key agreement rate and a 9–15× increase in key generation rate. Moreover, the proposed method attains a 4.37--9.33% improvement when adapted to new scenarios with limited data sizes. A security analysis demonstrates thatVehicle-Keyis resilient against several common attacks. Furthermore, we implementVehicle-Keyon a Raspberry Pi and demonstrate its ability to execute within 3.5 ms. Huanqi Yang, Di Duan, Hongbo Liu 0002, Chengwen Luo 0001, Yuezhong Wu, Wei Li 0058, Albert Y. Zomaya, Linqi Song, Weitao Xu |
IEEE Trans. Mob. Comput. | 2 |
| 2024 | mmSign: mmWave-based Few-Shot Online Handwritten Signature VerificationabstractHandwritten signature verification has become one of the most important document authentication methods that are widely used in the financial, legal, and administrative sectors. Compared with offline methods based on static signature images, online handwritten signature verification methods are more reliable because of the temporary dynamic information (e.g., signing velocity, writing force, stroke order) that alleviates the risk of being forged. However, most existing online handwritten signature verification solutions are reliant on specific signing devices (e.g., customized pens or writing pads) and require extensive data collection during the registration phase, resulting in poor adaptability and applicability for new users. In this article, we propose mmSign, a millimeter wave (mmWave)–based online handwritten signature verification system, which enables accurate sensing of the user’s hand movements when signing through the superior sensing capability of mmWave. mmSign extracts the time-velocity feature maps from the captured mmWave signals by the carefully designed signal processing algorithms and then exploits a transformer-based verification model for signature verification. In addition, a novel meta-learning strategy with proposed task generation and data augmentation methods is introduced in mmSign to teach the verification model to learn effectively with limited samples, allowing our model to quickly adapt to new users. Extensive experiments show that mmSign is a robust, efficient, and secure handwritten signature verification system, achieving 84.07%, 87.31%, 91.12%, and 96.54% verification accuracy when 1, 3, 5, and 10 labeled signatures are available, respectively, while being resistant to common forgery attacks. Mingda Han, Huanqi Yang, Tao Ni 0003, Di Duan, Mengzhe Ruan, Jia Zhang 0028, Weitao Xu |
ACM Trans. Sens. Networks | 4 |
| 2023 | EMGSense: A Low-Effort Self-Supervised Domain Adaptation Framework for EMG SensingabstractThis paper presents EMGSense, a low-effort self-supervised domain adaptation framework for sensing applications based on Electromyography (EMG). EMGSense addresses one of the fundamental challenges in EMG cross-user sensing—the significant performance degradation caused by time-varying biological heterogeneity—in a low-effort (data-efficient and label-free) manner. To alleviate the burden of data collection and avoid labor-intensive data annotation, we propose two EMG-specific data augmentation methods to simulate the EMG signals generated in various conditions and scope the exploration in label-free scenarios. We model combating biological heterogeneity-caused performance degradation as a multi-source domain adaptation problem that can learn from the diversity among source users to eliminate EMG heterogeneous biological features. To relearn the target-user-specific biological features from the unlabeled data, we integrate advanced self-supervised techniques into a carefully designed deep neural network (DNN) structure. The DNN structure can seamlessly perform two training stages that complement each other to adapt to a new user with satisfactory performance. Comprehensive evaluations on two sizable datasets collected from 13 participants indicate that EMGSense achieves an average accuracy of 91.9% and 81.2% in gesture recognition and activity recognition, respectively. EMGSense outperforms the state-of-the-art EMG-oriented domain adaptation approaches by 12.5%-17.4% and achieves a comparable performance with the one trained in a supervised learning manner. Di Duan, Huanqi Yang, Guohao Lan, Tianxing Li 0001, Xiaohua Jia, Weitao Xu |
PERCOM | 1 |