VLDB 2026 Research / reviewers in the wild / expert
Xinjun Cai
dblp:257/7719
· DBLP profile ↗
10ranked-venue papers
3as first author
10since 2021 · last 2024
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 6 · 2 first-author · 6 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | InNetScheduler: In-network scheduling for time- and event-triggered critical traffic in TSNabstractTime-Sensitive Networking (TSN) is an enabling technology for Industry 4.0. Traffic scheduling plays a key role for TSN to ensure low-latency and deterministic transmission of critical traffic. As industrial network scales, TSN networks are expected to support a rising number of both time-triggered and event-triggered critical traffic (TCT and ECT). In this work, we present InNetScheduler, the first in-network TSN scheduling paradigm that boosts the throughput, i.e., number of scheduled data flows, of both traffic types. Different from existing approaches that conduct entire scheduling on the server, InNetScheduler leverages the computation resources on switches to promptly schedule latency-critical ECT, and delegate the computational-intensive TCT scheduling to server. The key innovation of InNetScheduler includes a Load-Aware Optimizer to mitigate ECT conflicts, a Relaxated ECT Scheduler to accelerate in-network computation, and End-to-End Determinism Guarantee to lower scheduling jitter. We fully implement a suite of InNetScheduler-compatible TSN switches with hardwaresoftware co-design. Extensive experiments are conducted on both simulation and physical testbeds, and the results demonstrate InNetScheduler’s superior performance. By unleashing the power of in-network computation, InNetScheduler points out a direction to extend the capacity of existing industrial networks. Xiangwen Zhuge, Xinjun Cai, Xiaowu He, Zeyu Wang 0015, Fan Dang 0001, Zheng Yang 0002 |
INFOCOM | 2 |
| 2024 | TrinitySLAM: On-board Real-time Event-image Fusion SLAM System for DronesabstractDrones have witnessed extensive popularity among diverse smart applications, and visual Simultaneous Localization and Mapping (SLAM) technology is commonly used to estimate the six-degrees-of-freedom pose for drone flight control systems. However, traditional image-based SLAM cannot ensure the flight safety of drones, especially in challenging environments such as high-speed flight and high dynamic range scenarios. The event camera, a new vision sensor, holds the potential to enable drones to overcome these challenging scenarios if fused with the image-based SLAM. Unfortunately, the computational demands of event-image fusion SLAM have grown manifold compared with image-based SLAM. Existing research on visual SLAM acceleration cannot achieve real-time operation of event-image fusion SLAM on on-board computing platforms for drones. To fill this gap, we present TrinitySLAM , a high-accuracy, real-time, low-energy consumption event-image fusion SLAM acceleration framework utilizing Xilinx Zynq, an on-board heterogeneous computing platform. The key innovations of TrinitySLAM include a fine-grained computation allocation strategy, several novel hardware–software co-acceleration designs, and an efficient data exchange mechanism. We fully implement TrinitySLAM on the latest Zynq UltraScale+ platform and evaluate its performance on one custom-made drone dataset and four official datasets covering various scenarios. Comprehensive experiments show that TrinitySLAM improves the pose estimation accuracy by 28% with half end-to-end latency and 1.2× energy consumption reduction compared with the most comparable state-of-the-art heterogeneous computing platform acceleration baseline. Xinjun Cai, Jingao Xu, Kuntian Deng, Hongbo Lan, Yue Wu 0030, Xiangwen Zhuge, Zheng Yang 0002 |
ACM Trans. Sens. Networks | 1 |
| 2024 | Multi-User Mobile Augmented Reality with ID-Aware Visual InteractionabstractMost existing multi-user Augmented Reality (AR) systems only support multiple co-located users to view a common set of virtual objects but lack the ability to enable each user to directly interact with other users appearing in his/her view. Such multi-user AR systems should be able to detect the human keypoints and estimate device poses (for identifying different users) in the meantime. However, due to the stringent low latency requirements and the intensive computation of the preceding two capabilities, previous research only enables either of the two capabilities for mobile devices even with the aid of the edge server. Integrating the two capabilities is promising but non-trivial in terms of latency, accuracy, and matching. To fill this gap, we propose DiTing to achieve real-time ID-aware multi-device visual interaction for multi-user AR applications, which contains three key innovations: Shared On-device Tracking to merge the similar computation for optimized latency, Tightly Coupled Dual Pipeline to enhance the accuracy of each task through mutual assistance, and Body Affinity Particle Filter to precisely match device poses with human bodies. We implement DiTing on four types of mobile AR devices and develop a multi-user AR game as a case study. Extensive experiments show that DiTing can provide high-quality human keypoint detection and pose estimation in real time (30fps) for ID-aware multi-device interaction and outperform the state-of-the-art baseline approaches. Xinjun Cai, Zheng Yang 0002, Qiang Ma 0007 |
ACM Trans. Sens. Networks | 1 |
| 2023 | SHARK: A Lightweight Model Compression Approach for Large-scale Recommender SystemsabstractIncreasing the size of embedding layers has shown to be effective in improving the performance of recommendation models, yet gradually causing their sizes to exceed terabytes in industrial recommender systems, and hence the increase of computing and storage costs. To save resources while maintaining model performances, we propose SHARK, the model compression practice we have summarized in the recommender system of industrial scenarios. SHARK consists of two main components. First, we use the novel first-order component of Taylor expansion as importance scores to prune the number of embedding tables (feature fields). Second, we introduce a new row-wise quantization method to apply different quantization strategies to each embedding. We conduct extensive experiments on both public and industrial datasets, demonstrating that each component of our proposed SHARK framework outperforms previous approaches. We conduct A/B tests in multiple models on Kuaishou, such as short video, e-commerce, and advertising recommendation models. The results of the online A/B test showed SHARK can effectively reduce the memory footprint of the embedded layer. For the short-video scenarios, the compressed model without any performance drop significantly saves 70% storage and thousands of machines, improves 30% queries per second (QPS), and has been deployed to serve hundreds of millions of users and process tens of billions of requests every day. Beichuan Zhang 0002, Chenggen Sun, Jianchao Tan, Xinjun Cai, Mengqi Miao, Chengru Song, Na Mou, Yang Song 0008 |
CIKM | 4 |
| 2023 | HMSG: Heterogeneous graph neural network based on Metapath SubGraph learning
Mengya Guan, Xinjun Cai, Jiaxing Shang, Fei Hao 0001, Dajiang Liu, Xianlong Jiao, Wancheng Ni |
Knowl. Based Syst. | 2 |
| 2023 | WAVE: Edge-Device Cooperated Real-Time Object Detection for Open-Air ApplicationsabstractCNN based real-time object detection can facilitate various AI applications that need to understand the surroundings via camera, such as autonomous package delivery robots, augmented reality, and intelligent drone applications. Currently, due to the high computation cost of CNN, accurate real-time object detection is only possible when mobile devices can upload video frames to powerful edge servers through high-speed wireless networks like WiFi. However, for many open-air AI applications, the network conditions (such as cellular networks) are usually unfavorable, far from satisfying the network demands of state-of-the-art systems. In this paper, we focus on the challenges incurred by mobile communication networks and propose WAVEcontaining three novel techniques, which areDeep RoI Encoding,Prioritized Parallel OffloadingandFine-grained Offloading Strategy, to realizereal-time,robustandlow-costobject detection for open-air AI applications. The experimental results show that under LTE networks, WAVE realizes high-accuracy real-time object detection and face recognition and significantly outperforms state-of-the-art systems. Zheng Yang 0002, Xinjun Cai, Yi Zhao 0016, Qiang Ma 0007 |
IEEE Trans. Mob. Comput. | 3 |
| 2023 | Trine: Cloud-Edge-Device Cooperated Real-Time Video Analysis for Household ApplicationsabstractReal-time mobile video analysis like object detection and tracking is key to various household applications such as AR, cognitive assistance and smart home. Such applications rely on heavy DNN models, which are not suitable for mobile devices due to resource limitation. The long latency of cloud offloading is unacceptable for the real-time requirements, and the direct edge offloading relies on powerful edge servers, which is impractical for household scenarios. To solve this challenge, we take advantage of the computing devices that are low-cost or already exist in our lives, and propose Trine, a cloud-edge-device cooperated framework, in which complicated computation tasks are offloaded from the device to the cloud with the edge as the key bond to coordinate. In addition, due to the heterogeneity of edge devices, which leads to no one-fits-all algorithm that is optimal in all situations, we propose a profile-based algorithm to customize trackers for various edge devices. We implemented Trine on an android phone and three edge devices. The experiments demonstrate that Trine achieves 8-36% higher real-time accuracy and 25-89% higher robustness than state-of-the-art. Yi Zhao 0016, Zheng Yang 0002, Xiaowu He, Xinjun Cai, Qiang Ma 0007 |
IEEE Trans. Mob. Comput. | 4 |
| 2022 | DiTing: Edge Assisted Real-time ID-aware Visual Interaction for Multi-user Augmented RealityabstractMost existing multi-user Augmented Reality (AR) systems only support multiple co-located users to view a common set of virtual objects but lack the ability to enable each user to directly interact with other users appearing in his/her view. Such multi-user AR systems should be able to detect the human keypoints and share device poses (for identifying different users) in the meanwhile. However, due to the stringent low latency requirements and the intensive computation of the above two capabilities, previous research only enables either of the two capabilities for mobile devices even with the aid of the edge server. Integrating the above two capabilities is promising but non-trivial in terms of latency, accuracy, and matching. To fill this gap, we propose DiTing to achieve real-time ID-aware multi-device visual interaction for multi-user AR applications, which contains three key innovations: Shared On-device Tracking to merge the similar computation for optimized latency, Tightly Coupled Dual Pipeline to enhance the accuracy of each task through mutual assistance, Body Affinity Particle Filter to precisely match device poses with human bodies. We implement DiTing on four types of mobile AR devices and develop a multi-user AR game as a case study. Extensive experiments show that DiTing can provide high-quality human keypoint detection and pose estimation in real-time (30fps) for ID-aware multi-device interaction and outperform the SOTA baseline approaches. Xinjun Cai, Zheng Yang 0002, Qiang Ma 0007 |
ICPADS | 1 |
| 2021 | Gain Without Pain: Enabling Real-time Environmental Perception on 2x Mobile Devices in Multiplayer Augmented RealityabstractMobile multiplayer augmented reality(AR) emerges in various applications including games, education, training, etc. Edge computing technique enables real-time environmental perception ability(e.g. object detection and segmentation) of devices by offloading complex computation to the nearby edge server. However, with more players involved and offloading video streams, bandwidth competition intensifies and lengthens the transmission latency, which severely impairs the accuracy of environmental perception in multiplayer AR applications. We realize staggering offloading period of each device can reduce transmission latency and maintain real-time environmental perception when more players involved. We propose Bonus containing two key techniques: collaborative offloading scheduler to eliminate bandwidth competition among multiple devices, which improve the performance of the overall system to achieve “Pareto Efficiency”; time-aware video reshaper based on mainstream 802.11 protocol to enable Bonus compatible with most wireless scenarios. we evaluate Bonus and four SOTA solutions across 24 videos on different environmental perception tasks. Results demonstrate that Bonus achieves 119.5% accuracy and 1.9x player capacity compared to the closest baseline under wireless environment. Xinjun Cai, Zheng Yang 0002 |
ICPADS | 2 |
| 2021 | AIRCODE: Hidden Screen-Camera Communication on an Invisible and Inaudible Dual Channel
Kun Qian 0004, Yumeng Lu, Zheng Yang 0002, Kehong Huang, Xinjun Cai, Chenshu Wu, Yunhao Liu 0001 |
NSDI | 6 |