Sijie Ji

dblp:271/7071 · DBLP profile ↗
← Back
14ranked-venue papers
7as first author
13since 2021 · last 2026
0000-0002-6615-1982ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 7 · 6 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 3 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 AquaCode: Turning Transaction Platforms into Non-Invasive Liquid Quality Checker
abstract
Existing liquid sensing solutions typically fail to achieve real-world scalability. Current channel-based methods (e.g., RF, acoustics) approach this problem through a consumer-centric classification paradigm. However, their measurements become tightly coupled to non-repeatable channel responses and heterogeneous hardware. As a result, systems require massive, ad-hoc data collection to cover these variations, yet still fail to generalize across the practically infinite diversity of products. In this paper, we propose a paradigm shift to targeted verification and repurpose the identity tag as a sensing interface. We present AquaCode, the first liquid quality assessment system seamlessly integrated into retail checkout platforms. Instead of inferring liquid types via unstable channels, AquaCode harnesses Liquid-Solid Interface (LSI), a novel sensing modality that measures charge transfer dynamics directly at the outer container wall, making it immune to external multipath interference. Crucially, AquaCode leverages the checkout platform to provide standardized mechanical excitation, converting the erratic motion artifacts of handheld sensing into a deterministic signal source. This enables a robust verification model: the system compares the runtime LSI signal against a manufacturer-defined golden standard encoded in a low-cost passive tag. We design a specialized analog frontend to extract these faint, high-impedance signals and a lightweight algorithm for 1-to-1 verification. We achieve a recall of 0.905 for identifying low-quality drinks. We believe AquaCode offers a scalable path to ubiquitous liquid safety.
Sijie Ji, Sheng Lyu, Wei Gao 0004
MobiSys1
2026 MuSAC: Toward Communication-Free Sensory Data Acquisition
abstract
Sensing and communication are at the core of the Internet of Things, which usually function independently. For example, a smartphone can communicate over Wi-Fi or cellular networks while continuously acquiring sensory data from the environment through various sensors. This paper presents a novel framework, MuSAC (Mutualistic Sensing and Communication), which seamlessly integrates the collection of sensory data with existing communication systems, without adding any additional communication overhead (i.e., communication-free). The framework leverages the mutualistic relationship between specific communication data and sensory data to effectively crowdsource heterogeneous sensory data without harming communication performance in practical distributed systems. To embed massive sensory data into the current transmission of communication data, MuSAC presents novel neural networks to distill universal features from the raw data for compression at the sender side and then extract invariant features on the server side. By doing so, MuSAC eliminates additional communication overhead for sensory data collection while also mitigating privacy concerns and data heterogeneity in crowdsensing. To evaluate the performance of MuSAC, we first conduct system-level simulations in a distributed environment by embedding public human activity sensing datasets into cellular massive MIMO communication processes. The results demonstrate that MuSAC effectively supports heterogeneous sensory data collection over existing wireless communication links at zero additional communication overhead. We further implement a functional prototype of the MuSAC system using off-the-shelf devices, where real-world sensory data are seamlessly integrated into WiFi transmissions. This practical implementation demonstrates the viability of communication-free sensory data acquisition in real-world scenarios, paving the way for broader applications of MuSAC.
Sijie Ji, Lixiang Lian
IEEE Trans. Mob. Comput.1
2026 Neur IT: Pushing the Limit of Neural Inertial Tracking for Indoor Robotic IoT
abstract
Inertial tracking is vital for robotic IoT and has gained popularity thanks to the ubiquity of low-cost inertial measurement units and deep learning-powered tracking algorithms. Existing works, however, have not fully utilized IMU measurements, particularly magnetometers, nor have they maximized the potential of deep learning to achieve the desired accuracy. To address these limitations, we introduceNeurIT, which elevates tracking accuracy to a new level.NeurITemploys a Time-Frequency Block-recurrent Transformer (TF-BRT) at its core, combining both RNN and Transformer to learn representative features in both time and frequency domains. To fully utilize IMU information, we strategically employ body-frame differentiation of magnetometers, considerably reducing the tracking error. We implementNeurITon a customized robotic platform and conduct evaluation in various indoor environments. Experimental results demonstrate thatNeurITachieves a mere 1-meter tracking error over a 300-meter distance. Notably, it significantly outperforms state-of-the-art baselines by 48.21% on unseen data. Moreover,NeurITdemonstrates robustness in large urban complexes and performs comparably to the visual-inertial approach (Tango Phone) in vision-favored conditions while surpassing it in feature-sparse settings. We believeNeurITtakes an important step forward toward practical neural inertial tracking for ubiquitous and scalable tracking of robotic things.NeurITis open-sourced here:https://github.com/aiot-lab/NeurIT.
Xinzhe Zheng 0001, Sijie Ji, Yipeng Pan, Kaiwen Zhang 0016, Chenshu Wu
IEEE Trans. Mob. Comput.2
2025 Magnetometer-Calibrated Hybrid Transformer for Robust Inertial Tracking in Robotics
abstract
Inertial tracking is vital for autonomous robots and has gained popularity with the ubiquity of low-cost Inertial Measurement Units (IMUs) and deep learning-powered tracking algorithms. Existing works, however, have not fully utilized IMU measurements, particularly magnetometers, nor maximized the potential of deep learning to achieve the desired accuracy. To bridge the gap, we introduce NeurIT, which employs a Time-Frequency Block-recurrent Transformer (TF-BRT) at its core, combining RNN and Transformer to learn both time-frequency representative features. To fully utilize IMU information, we strategically employ differentiation of body-frame magnetometers for orientation calibration in a sensor fusion manner. Experiments conducted in diverse environments show that NeurIT maintains a mere 1 -meter tracking error over a 300 - meter distance, surpassing state-of-the-art baselines by 48.21 % on unseen data. NeurIT also performs comparably to the visual-inertial approach (Tango Phone) in vision-favored conditions and surpasses it in plain environments. We share the code and data to promote further research: https://github.com/aiot-lab/NeurIT.
Xinzhe Zheng 0001, Sijie Ji, Yipeng Pan, Kaiwen Zhang 0016, Jia Pan 0001, Chenshu Wu
ICRA2
2025 One-shot Federated Learning Methods: A Practical Guide
abstract
One-shot Federated Learning (OFL) is a distributed machine learning paradigm that constrains client-server communication to a single round, addressing privacy and communication overhead issues associated with multiple rounds of data exchange in traditional Federated Learning (FL). OFL demonstrates the practical potential for integration with future approaches that require collaborative training models, such as large language models (LLMs). However, current OFL methods face two major challenges: data heterogeneity and model heterogeneity, which result in subpar performance compared to conventional FL methods. Worse still, despite numerous studies addressing these limitations, a comprehensive summary is still lacking. To address these gaps, this paper presents a systematic analysis of the challenges faced by OFL and thoroughly reviews the current methods. We also offer an innovative categorization method and analyze the trade-offs of various techniques. Additionally, we discuss the most promising future directions and the technologies that should be integrated into the OFL field. This work aims to provide guidance and insights for future research.
Xiang Liu 0017, Zhenheng Tang, Xia Li 0005, Yijun Song, Sijie Ji, Bo Han 0003, Linshan Jiang, Jialin Li 0001
IJCAI5
2025 CardioLive: Empowering Video Streaming with Online Cardiac Monitoring via Audio-Visual Learning
Sheng Lyu, Ruiming Huang, Sijie Ji, Yasar Abbas Ur Rehman, Chenshu Wu
ACM Multimedia3
2025 GSRF: Complex-Valued 3D Gaussian Splatting for Efficient Radio-Frequency Data Synthesis
abstract
Synthesizing radio-frequency (RF) data given the transmitter and receiver positions, e.g., received signal strength indicator (RSSI), is critical for wireless networking and sensing applications, such as indoor localization. However, it remains challenging due to complex propagation interactions, including reflection, diffraction, and scattering. State-of-the-art neural radiance field (NeRF)-based methods achieve high-fidelity RF data synthesis but are limited by long training times and high inference latency. We introduce GSRF, a framework that extends 3D Gaussian Splatting (3DGS) from the optical domain to the RF domain, enabling efficient RF data synthesis. GSRF realizes this adaptation through three key innovations: First, it introduces complex-valued 3D Gaussians with a hybrid Fourier–Legendre basis to model directional and phase-dependent radiance. Second, it employs orthographic splatting for efficient ray–Gaussian intersection identification. Third, it incorporates a complex-valued ray tracing algorithm, executed on RF-customized CUDA kernels and grounded in wavefront propagation principles, to synthesize RF data in real time. Evaluated across various RF technologies, GSRF preserves high-fidelity RF data synthesis while achieving significant improvements in training efficiency, shorter training time, and reduced inference latency.
Kang Yang 0005, Gaofeng Dong, Sijie Ji, Wan Du, Mani Srivastava 0001
NeurIPS3
2025 Transforming Mental Health Care with Autonomous LLM Agents at the Edge
abstract
The integration of Large Language Models (LLMs) with mobile devices is set to transform mental health care accessibility and quality. This paper introduces MindGuard, an autonomous LLM agent that utilizes mobile sensor data and engages in proactive, personalized conversations while ensuring user privacy through local processing. Unlike traditional mental health AI tools, MindGuard enables real-time, context-aware interventions by dynamically adapting to users' emotional and physiological states. The real-world implementation demonstrates its effectiveness with the ultimate goal of creating an accessible, scalable, and personalized mental healthcare ecosystem for anyone with smart mobile devices.
Sijie Ji, Xinzhe Zheng 0001, Wei Gao 0004, Mani Srivastava 0001
SenSys1
2024 Predicting Adverse Events for Patients with Type-1 Diabetes Via Self-Supervised Learning
abstract
Predicting blood glucose levels is fundamental for precise primary care of type-1 diabetes (T1D) patients. However, it is challenging to predict glucose levels accurately, not to mention the early alarm of adverse events (hyperglycemia and hypoglycemia), namely the minority class. In this paper, we propose BG-BERT, a novel self-supervised learning framework for blood glucose level prediction. In particular, BG-BERT incorporates masked autoencoder to capture rich contextual information of blood glucose records for accurate prediction. More specifically, SMOTE data augmentation and shrinkage loss are employed to effectively handle adverse events without discrimination. We evaluate BG-BERT on two benchmark datasets against two state-of-the-art base-line models. The experimental results highlight the significant improvements achieved by BG-BERT in glucose level prediction accuracy (measured by RMSE) and sensitivity to adverse events, with average lifting ratios of 9.5% and 44.9%, respectively.
Xinzhe Zheng 0001, Sijie Ji, Chenshu Wu
ICASSP2
2024 MuSAC: Mutualistic Sensing and Communication for Mobile Crowdsensing
abstract
Sensing and communication are at the core of the Internet of Things, which usually function independently. For example, a smartphone can communicate over Wi-Fi or cellular networks while continuously acquiring sensory data from the environment through various sensors. This paper presents a novel framework, MuSAC (Mutualistic Sensing and Commu-nication), which seamlessly integrates the collection of sensory data with existing communication systems, without adding any extra communication overhead. The framework leverages the mutualistic relationship between specific communication data and sensory data to effectively crowdsource heterogeneous sensory data without harming communication performance in practical distributed systems. To embed massive sensory data into the current transmission of communication data, MuSAC presents novel neural networks to distill universal features from the raw data for compression at the sender side and then extract invariant features on the server side. By doing so, MuSAC eliminates additional communication costs for sensory data collection while also mitigating privacy concerns and data heterogeneity in crowd-sensing. Our real-world experimental validation in Wi-Fi and cellular Massive MIMO communication scenarios demonstrates the effectiveness of the MuSAC framework, shedding light on efficient mobile crowdsensing for massive IoT data collection.
Sijie Ji, Lixiang Lian, Yuanqing Zheng, Chenshu Wu
ICDCS1
2023 Enhancing Deep Learning Performance of Massive MIMO CSI Feedback
abstract
CSI feedback is an important problem of massive multiple-input multiple-output (MIMO) technology because the feedback overhead is proportional to the number of sub-channels and the number of antennas, both of which scale with the size of the massive MIMO system. Deep learning-based CSI feedback methods have been widely adopted recently owing to their superior performance. Despite the success, current approaches have not fully exploited the relationship between the characteristics of CSI data and the deep learning framework. In this paper, we propose a jigsaw puzzles aided training strategy (JPTS) to enhance the deep learning-based massive MIMO CSI feedback approaches by maximizing mutual information between the original CSI and the compressed CSI. We apply JPTS on top of existing state-of-the-art methods. Experimental results show that by adopting this training strategy, the accuracy can be boosted by 12.07% and 7.01% on average in indoor and outdoor environments, respectively. The proposed method is ready to adopt to any existing deep learning frameworks of massive MIMO CSI feedback. Codes of JPTS are available on GitHub for use11https://github.com/SIJIEJI/JPTS.
Sijie Ji, Mo Li 0001
ICC1
2023 Construct 3D Hand Skeleton with Commercial WiFi
abstract
This paper presents HandFi, which constructs hand skeletons with practical WiFi devices. Unlike previous WiFi hand sensing systems that primarily employ predefined gestures for pattern matching, by constructing the hand skeleton, HandFi can enable a variety of downstream WiFi-based hand sensing applications in gaming, healthcare, and smart homes. Deriving the skeleton from WiFi signals is challenging, especially because the palm is a dominant reflector compared with fingers. HandFi develops a novel multi-task learning neural network with a series of customized loss functions to capture the low-level hand information from WiFi signals. During offline training, HandFi takes raw WiFi signals as input and uses the leap motion to provide supervision. During online use, only with commercial WiFi, HandFi is capable of producing 2D hand masks as well as 3D hand poses. We demonstrate that HandFi can serve as a foundation model to enable developers to build various applications such as finger tracking and sign language recognition, and outperform existing WiFi-based solutions. Artifacts can be found: https://github.com/SIJIEJI/HandFi
Sijie Ji, Xuanye Zhang, Yuanqing Zheng, Mo Li 0001
SenSys1
2022 SiFall: Practical Online Fall Detection with RF Sensing
abstract
Falls are one of the leading causes of death in the elderly people aged 65 and above. In order to prevent death by sending prompt fall detection alarms, non-invasive radio-frequency (RF) based fall detection has attracted significant attention, due to its wide coverage and privacy preserving nature. Existing RF-based fall detection systems process fall as an activity classification problem and assume that human falls introduce reproducible patterns to the RF signals. We, however, argue that the fall is essentially an accident, hence, its impact is uncontrollable and unforeseeable. We propose to solve the fall detection problem in a fundamentally different manner. Instead of directly identifying the human falls which are difficult to quantify, we recognize the normal repeatable human activities and then identify the fall as abnormal activities out of the normal activity distribution. We implement our idea and build a prototype based on commercial Wi-Fi. We conduct extensive experiments with 16 human subjects. The experiment results show that our system can achieve high fall detection accuracy and adapt to different environments for real-time fall detection.
Sijie Ji, Yaxiong Xie, Mo Li 0001
SenSys1
2020 Attention-driven Unsupervised Image Retrieval for Beauty Products with Visual and Textual Clues
abstract
Beauty and personal care product retrieval (BPCR) aims to match a query image of an item to examples of the same item in a large database. The task is extremely challenging because a small number of ground-truth examples have to be found in a large search space. Previous works mostly search only with visual representations and have not made full use of the product descriptions. Since many noisy examples only have subtle visual differences comparing to the ground-truth examples (e.g. similar packaging but different brands) and those differences (e.g. product brands) are especially hard to be captured only by visual features, methods merely based on visual feature similarities can easily regard those noisy examples as examples of the same item in the query image. We notice that the product descriptions are good sources for capturing those subtle visual differences. Therefore, we propose a search method utilizing both images and product descriptions in this work. Before searching, we not only prepare attention-based visual features for each database image but also a textual index (TI) that matches each database example to other examples with similar product descriptions. During searching, the visual feature of the query image is firstly searched in the whole database and then searched in a subset obtained by looking up the TI. Finally, the second result is used to refine the initial result. Since the subset examples usually have similar properties (e.g. brands and type), the noisy examples in the initial result can be effectively replaced. We have experimentally proved the effectiveness of the proposed method on the validation set of the Perfect-500K dataset. Our team (NTU-Beauty) achieved the 3rd place in the leader board of the Grand Challenge of AI Meets Beauty in ACM Multimedia 2020. Our code is available at: https://github.com/jingwenh/2020-ai-meets-beauty_ntubeauty.git.
Jingwen Hou, Sijie Ji, Annan Wang
ACM Multimedia2