VLDB 2026 Research / reviewers in the wild / expert
Bojun Zhang 0001
dblp:247/5223-1
· DBLP profile ↗
6ranked-venue papers
1as first author
6since 2021 · last 2026
0009-0008-8159-5822ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 6 · 1 first-author · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Physics-Aware Multifeature Fusion Approach for Robust Channel EstimationabstractAccurate CSI feedback is crucial for Massive MIMO systems, yet it remains challenging in resource-constrained IoT scenarios due to strict pilot overhead constraints. Under such extreme data sparsity, conventional data-driven methods often fail to generalize. To address this, this paper proposes a Physics-Aware Multi-Feature fusion approach (PAMF), a deep learning framework that systematically integrates data-driven learning with wireless propagation physics. The framework includes dedicated feature extractors based on spatial, frequency-domain, statistical, and physics-based methods, along with a deep residual reconstruction network. A key innovation lies in its dual-level physical constraint mechanism, which incorporates domain knowledge at both the feature and loss levels to ensure physically plausible channel estimates. By leveraging multi-modal feature representations and physics-aware optimization, PAMF effectively recovers the complete channel matrix from sparse pilot signals, which not only improves feature discrimination, but also leads to greater robustness particularly under the dynamic conditions typical of urban mobile networks. Experimental results demonstrate that the proposed method consistently outperforms existing approaches across diverse datasets including MIMO configurations of various scales, different modulation schemes, and real-world Wi-Fi CSI Specifically, on real-world Wi-Fi data, PAMF achieves an NMSE of 0.1002, approximately 62% lower than the ChannelNet baseline (0.2669). Overall, this study contributes a practical and physical-aware framework for channel estimation, paving the way for more reliability and efficiency next-generation wireless systems, with direct implications for large-scale IoT deployments. Jiancheng Chen, Jiuwu Zhang, Bojun Zhang 0001, Xiaomin Zhou, Mingli Feng, Keqiu Li |
IEEE Internet Things J. | 3 |
| 2024 | A Wireless Signal Correlation Learning Framework for Accurate and Robust Multi-Modal SensingabstractWireless signal analytics in IoT systems can enable various promising wireless sensing applications such as localization, anomaly detection, and human activity recognition. As a matter of fact, there are significant correlations in terms of dimension, spatial and temporal aspects among wireless signals from multiple sensors. However, none of the wireless sensing research currently in use directly incorporates or exploits the signal correlations. Therefore, there is still substantial scope for improvement in regards to accuracy and robustness. We are introducing a novel framework called Signal Correlation Learning (SCL). This framework utilizes a directed graph to explicitly represent the signal correlation across various wireless sensors. We use signal embedding to depict the correlation features of a multi-dimensional sensor that arise from a multi-sensor system. Then, we perform Kullback-Leibler (KL) divergence on embedding vectors of any pair of sensors in the system to construct a subgraph at a given time point, which can measure the spatial signal correlation of sensors. Subsequently, several subgraphs spanning a specific time frame are fused into a coherent universal graph based on the small-world theory. This universal graph represents the three types of signal correlation simultaneously. A signal correlation aggregation structure is utilized to extract the features from the universal graph. These features can be used to address target sensing problems. We implement SCL in real RFID, Bluetooth, WIFI, and Zigbee systems, and evaluate its performance in three common wireless sensing problems including localization, anomaly detection, and human activity recognition. Extensive experiments demonstrate that our SCL framework significantly outperforms state-of-the-art wireless sensing algorithms by increasing$80\%\sim 190\%$in terms of accuracy, and by increasing$160\%\sim 220\%$in terms of robustness. Xiulong Liu 0001, Bojun Zhang 0001, Sheng Chen 0015, Xin Xie 0001, Xinyu Tong 0001, Tao Gu 0001, Keqiu Li |
IEEE J. Sel. Areas Commun. | 2 |
| 2024 | Fine-Grained Recognition of Manipulation Activities on Objects via Multi-Modal SensingabstractFine-grained recognition of human manipulation activities on objects is crucial in the era of human-computer-object integration. However, there is a lack of solutions for simultaneous recognition of human identity, manipulation activities (including drawing and rotation), and manipulated objects. Therefore, we propose an RF-Camera system that combines RFID and computer vision techniques to address this challenge in multi-person and multi-object scenarios. In RF-Camera, we employ a skeleton-assisted method to extract facial images of target individuals, enabling precise recognition of their identities. To identify manipulation activities, we analyze the 3D hand trajectory and fingertip vector angle, differentiating drawing and rotation manipulation activities. Additionally, we model target person?s hand movements to predict phase data of the target tag, enabling the determination of person-object relationships. Implementing RF-Camera using COTS RFID and Kinect devices involves overcoming challenges such as extracting effective data from noisy streams, predicting virtual phase data considering hand-tag offset, and ensuring high tag reading rates in tag-dense scenarios. We conducted experiments involving six participants performing object manipulation activities, including drawing letters/symbols and rotating movements. Extensive experimental results show that RF-Camera achieves over 90% accuracy in recognizing person identity, manipulation activities, and person-object matching in most conditions. Xiulong Liu 0001, Bojun Zhang 0001, Lizhang Wang, Sheng Chen 0015, Xin Xie 0001, Xinyu Tong 0001, Tao Gu 0001, Keqiu Li |
IEEE Trans. Mob. Comput. | 2 |
| 2022 | An RFID and Computer Vision Fusion System for Book Inventory using Mobile RobotabstractMobile robot-assisted book inventory such as book identification and book order detection has become increasingly popular in smart library, replacing the manual book inventory which is time-consuming and error-prone. The existing systems are either computer vision (CV)-based or RFID-based, however several limitations are inevitable. CV-based systems may not be able to identify books effectively due to low accuracy of detecting texts on book spine. RFID tags attached to books can be used to identify a book uniquely. However, in high tag density scenarios such as library, tag coupling effects of adjacent tags may seriously affect the accuracy of tag reading. To overcome these limitations, this paper presents a novel RFID and CV fusion system for Book Inventory using mobile robot (RC-BI). RFID and CV are first used individually to obtain book order, then the information will be fused by the sequence based matching algorithm to remove ambiguity and improve overall accuracy. Specifically, we address three technical challenges. We design a deep neural network (DNN) model with multiple inputs and mixed data to filter out interference of RFID tags on other tiers, and propose a video information extracting schema to extract book spine information accurately, and use strong link to align and match RFID- and CV-based timestamp vs. book-name sequences to avoid errors during fusion. Extensive experiments indicate that our system achieves an average accuracy of 98.4% for tier filtering and an average accuracy of 98.9% for book order, significantly outperforming the state-of-the-arts. Jiuwu Zhang, Xiulong Liu 0001, Tao Gu 0001, Bojun Zhang 0001, Zijuan Liu, Keqiu Li |
INFOCOM | 4 |
| 2022 | RC6D: An RFID and CV Fusion System for Real-time 6D Object Pose EstimationabstractThis paper studies the problem of 6D pose estimation, which is practically important in various application scenarios such as robotic-based object grasping, obstacle avoidance in autonomous driving scene, and object integration in mixed reality. However, existing methods suffer from at least one of the five major limitations: dependence on object identification, complex deployment, difficulty in data collection, low accuracy, and incomplete estimation. To overcome the above limitations, this paper proposes an RC6D system, which is the first to estimate 6D poses by fusing RFID and Computer Vision (CV) data with multi-modal deep learning techniques. In RC6D, we first detect 2D keypoints through a deep learning approach. We then propose a novel RFID-CV fusion neural network to predict the depth of the scene, and use the estimated depth information to expand the 2D keypoints to 3D keypoints. Finally, we model the coordinate correspondences between the detected 2D-3D keypoints, which is applied to estimate the 6D pose of the target object. When implementing RC6D, we mainly address the following three technical challenges. (i) To predict 6D poses without using the CAD model, we propose a network architecture for monocular depth estimation. (ii) To train the neural network for 6D pose estimation without time-consuming 6D labeling, we use an unsupervised learning algorithm based on 2D-3D point pair matching. (iii) To detect the subject of the object without identification, we leverage optical flow to restrict the object and RFID to directly obtain its information. The experimental results show that the localization error of RC6D is less than 10 cm with a probability higher than 90.64% and its orientation estimation error is less than 10° with a probability higher than 79.63%. Hence, the proposed RC6D system performs much better than the state-of-the-art related solutions. Bojun Zhang 0001, Mengning Li, Xin Xie 0001, Luoyi Fu, Xinyu Tong 0001, Xiulong Liu 0001 |
INFOCOM | 1 |
| 2021 | A Lightweight Heatmap-based Eye Tracking SystemabstractEye tracking is playing an important role in many applications including human-computer interaction and behavior study. However, the existing approaches have at least one of the following limitations: (i) dedicated devices such as infrared camera and eye-tracker are required; (ii) complex calibration process is involved; (iii) substantial computing resources are consumed; (iv) users suffer from the risk of privacy leakage. To address the above limitations, we propose a H eatmap-based E ye T racking (HETrack) system. One of the key challenges in our system is to design a lightweight model for fine-grained tracking when the computing resources of device is limited. Also, it is necessary to protect user privacy in such a system. To address the above challenging issues, the proposed system consists of the following processes. First, when users randomly look at the screen of the device, HETrack obtains the raw image containing facial information. Then, we design a neural network model and train it with federated learning. The model can map the image to heatmap that implies the possibility of the user’s gaze position on the screen. Finally, HETrack can intercept the real-time video stream into frames, and employ the trained model to generate the heatmap of current frame for gaze estimation. We implement HETrack based on a Commercial-Off-The-Shelf (COTS) camera and conduct extensive experiments to evaluate its performance. Our HETrack system only requires once calibration; whereas, the state-of-the-art work proposed by Google requires 3~5 times calibration on average. Unlike previous approaches that transmit raw image data to a central server, in our HETrack system, only parameters are transmitted, thereby well protecting the user’s privacy. Experimental results demonstrate that the average distance error of estimated gaze point is 3cm, which is compatible with the state-of-the-art methods. Xiaoxiao Luan, Bojun Zhang 0001, Xiulong Liu 0001, Xinyu Tong 0001, Keqiu Li |
ICCCN | 2 |