EDBT 2026 Demo / reviewers in the wild / expert
Xuewei Bai
dblp:63/10177
· DBLP profile ↗
8ranked-venue papers
3as first author
8since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 3 · 2 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
5 papers |
Robot navigation and mapping · 41% Video understanding and tracking · 30% 3D vision · 16% | |
| Computer networks
1 paper |
Vehicular, aerial and satellite networks · 50% Internet of things and sensor networks · 50% | |
| Theoretical computer science
2 papers |
Graph algorithms and graph theory · 100% |
Topics — the 16 heaviest of 16, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Video understanding and tracking
multi-object tracking |
1.6 | 2 | 2025 | STAR: Spatial-Temporal Tracklet Matching for Multi-Object Tracking · NeurIPS 2025 GSLAMOT: A Tracklet and Query Graph-based Simultaneous Locating, Mapping, and Multiple Object Tracking System · ACM Multimedia 2024 |
Robotics › Robot navigation and mapping
SLAM |
1.4 | 2 | 2024 | GSLAMOT: A Tracklet and Query Graph-based Simultaneous Locating, Mapping, and Multiple Object Tracking System · ACM Multimedia 2024 ColSLAM: A Versatile Collaborative SLAM System for Mobile Phones Using Point-Line Features and Map Caching · ACM Multimedia 2023 |
Computer vision › Video understanding and tracking › multi-object tracking
tracklet association |
0.9 | 1 | 2025 | STAR: Spatial-Temporal Tracklet Matching for Multi-Object Tracking · NeurIPS 2025 |
Graph algorithms and graph theory
graph matching |
0.9 | 1 | 2025 | STAR: Spatial-Temporal Tracklet Matching for Multi-Object Tracking · NeurIPS 2025 |
Knowledge, reasoning and agents › Multi-agent systems
distributed estimation |
0.7 | 1 | 2023 | Communication Efficient, Distributed Relative State Estimation in UAV Networks · IEEE J. Sel. Areas Commun. 2023 |
Robotics › Robot navigation and mapping › SLAM
multi-robot SLAM |
0.7 | 1 | 2023 | ColSLAM: A Versatile Collaborative SLAM System for Mobile Phones Using Point-Line Features and Map Caching · ACM Multimedia 2023 |
Computer vision › 3D vision
point cloud analysis |
0.7 | 1 | 2023 | ViPFormer: Efficient Vision-and-Pointcloud Transformer for Unsupervised Pointcloud Understanding · ICRA 2023 |
Robotics › Robot navigation and mapping › state estimation
relative state estimation |
0.7 | 1 | 2023 | Communication Efficient, Distributed Relative State Estimation in UAV Networks · IEEE J. Sel. Areas Commun. 2023 |
Computer vision › 3D vision › point cloud analysis › point cloud learning
unsupervised point cloud learning |
0.7 | 1 | 2023 | ViPFormer: Efficient Vision-and-Pointcloud Transformer for Unsupervised Pointcloud Understanding · ICRA 2023 |
Robotics › Robot navigation and mapping › visual odometry
visual-inertial odometry |
0.7 | 1 | 2023 | ColSLAM: A Versatile Collaborative SLAM System for Mobile Phones Using Point-Line Features and Map Caching · ACM Multimedia 2023 |
Internet of things and sensor networks › wireless sensor network › distributed algorithms for sensor networks
distributed state estimation |
0.7 | 1 | 2023 | Communication Efficient, Distributed Relative State Estimation in UAV Networks · IEEE J. Sel. Areas Commun. 2023 |
Vehicular, aerial and satellite networks › aerial networks
UAV networks |
0.7 | 1 | 2023 | Communication Efficient, Distributed Relative State Estimation in UAV Networks · IEEE J. Sel. Areas Commun. 2023 |
Machine learning › Representation and self-supervised learning
contrastive learning |
0.2 | 1 | 2023 | ViPFormer: Efficient Vision-and-Pointcloud Transformer for Unsupervised Pointcloud Understanding · ICRA 2023 |
Machine learning › Representation and self-supervised learning › contrastive learning
multimodal contrastive learning |
0.2 | 1 | 2023 | ViPFormer: Efficient Vision-and-Pointcloud Transformer for Unsupervised Pointcloud Understanding · ICRA 2023 |
Virtual and augmented reality › augmented reality
mobile augmented reality |
0.2 | 1 | 2023 | ColSLAM: A Versatile Collaborative SLAM System for Mobile Phones Using Point-Line Features and Map Caching · ACM Multimedia 2023 |
Graph algorithms and graph theory
graph optimization |
0.2 | 1 | 2023 | Communication Efficient, Distributed Relative State Estimation in UAV Networks · IEEE J. Sel. Areas Commun. 2023 |
Methods — techniques the papers use, named apart from their topics
message propagation · 1.7graph neural network · 1.7map caching · 1.3graph optimization · 1.3distributed graph optimization · 1.3RIPPLE-like iteration · 1.3NetVLAD · 1.3tracklet graph · 0.8query graph · 0.8vision transformer · 0.7point-line fusion · 0.7cross-modal learning · 0.7contrastive learning · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | STAR: Spatial-Temporal Tracklet Matching for Multi-Object TrackingabstractExisting tracking-by-detection Multi-Object Tracking methods mainly rely on associating objects with tracklets using motion and appearance features. However, variations in viewpoint and occlusions can result in discrepancies between the features of current objects and those of historical tracklets. To tackle these challenges, this paper proposes a novel Spatial-Temporal Tracklet Graph Matching paradigm (STAR). The core idea of STAR is to achieve long-term, reliable object association through the association of ``tracklet clips (TCs)". TCs are segments of confidently associated multi-object trajectories, which are linked through graph matching. Specifically, STAR initializes TCs using a Confident Initial Tracklet Generator (CITG) and constructs a TC graph via Tracklet Clip Graph Construction (TCGC). In TCGC, each object in a TC is treated as a vertex, with the appearance and local topology features encoded on the vertex. The vertices and edges of the TC graph are then updated through message propagation to capture higher-order features. Finally, a Tracklet Clip Graph Matching (TCGM) method is proposed to efficiently and accurately associate the TCs through graph matching. STAR is model-agnostic, allowing for seamless integration with existing methods to enhance their performance. Extensive experiments on diverse datasets, including MOTChallenge, DanceTrack, and VisDrone2021-MOT, demonstrate the robustness and versatility of STAR, significantly improving tracking performance under challenging conditions. Xuewei Bai, Yongcai Wang, Deying Li 0001, Haodi Ping, Chunxu Li |
NeurIPS | 1 |
| 2025 | SGB-YOLOv5: straw granulator blockage monitoring system
Haoyang Tong, Dongyang Gao, Zhixu Wang, Longlong Feng, Xuewei Bai |
J. Supercomput. | 6 |
| 2025 | A Geometric and Hypothesis-Based Method for Low-Overlap, Sparse, and Featureless Point Set MatchingabstractThis article proposes a general solution for point set matching that effectively addresses the challenges of low-overlap, sparse, or featureless point set matching (LSFPM). Unlike previous methods that mainly rely on feature or neighborhood similarity that often fail under such difficult conditions, this work proposes a Geometry-based Point Matching (GPM) method. GPM first introduces two geometric concepts: the “Structural Element” (SE) and the “Superstructural Element” (SSE), both of which are constructed based on local geometric structures. The SSE is an enhanced version of the SE. A descriptor for the SE, called the SE Descriptor (SED), is designed to encode the SE and facilitate an efficient geometry-based similarity metric. We demonstrate that the cosine similarity of SEDs is invariant to scale and rotation. Subsequently, a SE Matching Maximization (SEMM) problem is formulated to identify a size-penalized SSE set that maximizes the sum of similarities. This problem is efficiently solved using the proposed SEMM-MCMC (Markov Chain Monte Carlo) algorithm. The matched SSEs then vote on corresponding point matches, generating high-confidence one-to-one matches, low-confidence one-to-one matches and one-to-many matches. Finally, the InferMatch algorithm is proposed to jointly assess low-confidence one-to-one point set matching while simultaneously distinguishing one-to-many point set matching. The GPM approach can also complement other feature-based and motion-based methods. It has been extensively validated on both synthetic and real datasets, demonstrating its versatility in addressing various point set matching problems, and is not limited to the LSFPM problem. Extensive experiments on diverse datasets, including SPair-71k, UAVDT, VisDrone2021-MOT, and SparseMatch, further demonstrate the robustness and versatility of GPM. The proposed approach significantly improves matching performance under challenging conditions and effectively addresses key limitations of existing point set matching methods. Xuewei Bai, Yongcai Wang, Peng Wang 0106, Chunxu Li, Shuo Wang 0015, Deying Li 0001 |
ACM Trans. Sens. Networks | 1 |
| 2024 | GSLAMOT: A Tracklet and Query Graph-based Simultaneous Locating, Mapping, and Multiple Object Tracking System
Shuo Wang 0015, Yongcai Wang, Yongyu Guo, Xuewei Bai, Deying Li 0001 |
ACM Multimedia | 7 |
| 2024 | InferLoc: Hypothesis-Based Joint Edge Inference and Localization in Sparse Sensor NetworksabstractRanging-based localization is a fundamental problem in the Internet of Things and unmanned aerial vehicle networks. However, the nodes’ limited-ranging scope and users’ broad coverage purpose inevitably cause network sparsity or subnetwork sparsity. The performances of existing localization algorithms are extremely unsatisfactory in sparse networks. A crucial way to deal with the sparsity is to exploit the hidden knowledge provided by the unmeasured edges, which inspires this work to propose a hypothesis-based Joint Edge Inference and Localization algorithm called InferLoc . InferLoc mines the Unmeasured but Inferable Edges (UIEs). Each UIE is an unmeasured edge, but it is restricted through other edges in the network to be inside a rigid component, so it has only a limited number of possible lengths. We propose an efficient method to detect UIEs and geometric approaches to infer possible lengths for UIEs in 2D and 3D networks. The inferred possible lengths of UIEs are then treated as multiple hypotheses to determine the node locations and the lengths of UIEs simultaneously through a joint graph optimization process. In the joint graph optimization model, to make the 0/1 decision variables for hypotheses selection differentiable, differentiable functions are proposed to relax the 0/1 selections, and rounding is applied to select the final length after the optimization converges. We also prove the condition when a UIE can contribute to sparse localization. Extensive experiments show remarkably better accuracy and efficiency performances of InferLoc than the state-of-the-art network localization algorithms. In particular, it reduces the localization errors by more than 90% and speeds up the convergence time more than 100 times than that of the widely used G2O-based methods in sparse networks. Xuewei Bai, Yongcai Wang, Haodi Ping, Xiaojia Xu, Deying Li 0001, Shuo Wang 0015 |
ACM Trans. Sens. Networks | 1 |
| 2023 | ViPFormer: Efficient Vision-and-Pointcloud Transformer for Unsupervised Pointcloud UnderstandingabstractRecently, a growing number of work design unsupervised paradigms for point cloud processing to alleviate the limitation of expensive manual annotation and poor transferability of supervised methods. Among them, CrossPoint follows the contrastive learning framework and exploits image and point cloud data for unsupervised point cloud understanding. Although the promising performance is presented, the unbalanced architecture makes it unnecessarily complex and inefficient. For example, the image branch in CrossPoint is ~8.3x heavier than the point cloud branch leading to higher complexity and latency. To address this problem, in this paper, we propose a lightweight Vision-and-Pointcloud Transformer (ViPFormer) to unify image and point cloud processing in a single architecture. ViPFormer learns in an unsupervised manner by optimizing intra-modal and cross-modal contrastive objectives. Then the pretrained model is transferred to various downstream tasks, including 3D shape classification and semantic segmentation. Experiments on different datasets show ViPFormer surpasses previous state-of-the-art unsupervised methods with higher accuracy, lower model complexity and runtime latency. Finally, the effectiveness of each component in ViPFormer is validated by extensive ablation studies. The implementation of the proposed method is available at https://github.com/auniquesun/ViPFormer. Hongyu Sun 0006, Yongcai Wang, Xuewei Bai, Deying Li 0001 |
ICRA | 4 |
| 2023 | ColSLAM: A Versatile Collaborative SLAM System for Mobile Phones Using Point-Line Features and Map CachingabstractOver the past years, augmented reality (AR) based on mobile phones has gained great attention. When multiple phones are used in AR applications, collaborative simultaneous localization and mapping (SLAM) is considered one of the enabling technologies, i.e., multiple mobile phones complete the localization and mapping through collaboration. However, the state-of-the-art collaborative SLAM systems not only suffer from the delays introduced by a high-complexity graph optimization problem, but also may exhibit varying levels of accuracy across dissimilar environments or different types of mobile devices. In this paper, we propose a scalable and robust collaborative SLAM system, point-line-based Collaborative SLAM (ColSLAM). Technically, ColSLAM includes two innovative features that help achieve satisfactory scalability and robustness. First, a mapping cacher (MC) is designed for each agent on the server, which uses global keyframes to detect loop closures, updates the cached local map, and quickly responds to the agent's pose drifts. With MC, each agent's local pose is corrected using global knowledge in real-time. Secondly, to improve the robustness performance, ColSLAM employs point-line-fusion-based Visual Inertial Odometry (VIO), point-line-fusion-based NetVLAD loop detection, and an enhanced geometric verification and relative pose calculation method called PNPL. Empirical evaluations based on the EuRoc dataset and real degenerate environments demonstrate that ColSLAM outperforms the existing collaborative SLAM systems in terms of accuracy, robustness, and scalability. Yongcai Wang, Yongyu Guo, Shuo Wang 0015, Xuewei Bai, Qiang Ye 0001, Deying Li 0001 |
ACM Multimedia | 6 |
| 2023 | Communication Efficient, Distributed Relative State Estimation in UAV NetworksabstractDistributed estimation of 6-DOF relative states, including three-dimensional relative poses and three-dimensional relative positions, is a key problem in UAV (Unmanned Aerial Vehicle) networks, which generally requires vision-involved iterative state estimation. How to achieve communication efficiency is a crucial challenge considering the large volume of vision data. This paper jointly considers the communication efficiency, latency, and accuracy for distributed relative state estimation involving vision data in UAV networks. The key is to solve a distributed graph optimization problem, which includes two key steps: (1) local graph construction and node state initialization in an initialization phase, and (2) iterative state update and communication with neighbors until convergence in online iteration phase. A communication efficient, Locating Then Informing (LTI) initialization scheme is proposed, which is run only once by each node to initialize each node’s local graph and initial states. For online iteration, a RIPPLE-like distributed state iteration scheme is proposed. It inherits the advantages of traditional sequential and parallel methods while avoiding their drawbacks. It enables nodes’ states to converge quickly using fewer rounds of communications. The communication costs for the initialization and online iteration processes are analyzed theoretically. Extensive evaluations use synthetic data generated by AirSim (a widely used UAV network simulation platform) and real-world data are presented. The results show that the proposed method provides accuracy comparable to the centralized graph optimization method and significantly outperforms the other distributed methods in terms of accuracy, communication cost, and latency. Shuo Wang 0015, Yongcai Wang, Xuewei Bai, Deying Li 0001 |
IEEE J. Sel. Areas Commun. | 3 |