Kaiqi Chen 0001

dblp:273/7157-1 · DBLP profile ↗
← Back
10ranked-venue papers
2as first author
10since 2021 · last 2025
0000-0002-9079-0899ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 5 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Systems, architecture and hardware · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Alice-SLAM: Accurate and Lite-Communication Collaborative SLAM for Resource-Constrained Multi-Agent
abstract
Multi-agent collaborative simultaneous localization and mapping (Mac-SLAM) facilitates mutual localization among multi-agent and mapping in unknown environments. However, Mac-SLAM faces two main practical challenges in resource-constrained situations: heavy communication load and conflicts among multi-source maps. To address these issues, we propose Alice-SLAM: an accurate and lite-communication client-server collaborative SLAM system, reducing communication load while accuracy-guaranteed. Specifically, regarding high communication demand, we optimize communication load by compressing keyframe data and sharing only key map information instead of full map information. For inconsistency among multi-maps, we combine specific bundle adjustments (BA) and an adaptive strategy for active map optimization to enhance the consistency of the global map. A set of experiments demonstrates the superior accuracy and reduced communication load of the proposed Alice-SLAM on the EuRoC dataset and in multi-user augmented reality (AR) experiments conducted in our lab, highlighting its effectiveness in resource-constrained cases. We plan to open-source our code1to encourage further research and collaboration in this area.
Kaiqi Chen 0001, Ruyu Liu, Xu Cheng 0003, Jianhua Zhang 0002, Shengyong Chen, Houxiang Zhang, Arash Ajoudani
IEEE J. Sel. Areas Commun.2
2025 Semantic Visual Simultaneous Localization and Mapping: A Survey
abstract
Visual Simultaneous Localization and Mapping (vSLAM) is a cornerstone technology in computer vision and robotics, underpinning applications such as autonomous vehicles and robot navigation. While traditional vSLAM systems have shown significant progress in indoor or outdoor environments, their performance often degrades in complex scenes, limiting their adaptability and robustness. Semantic vSLAM, which integrates high-level semantic information into vSLAM systems, has emerged as a promising solution to address these limitations by enabling a richer understanding of the environment. In this paper, we provide a comprehensive review of semantic vSLAM, offering a critical analysis of its evolution, methods, and challenges. We begin by revisiting the development of traditional vSLAM, emphasizing its limitations and the motivation for incorporating semantic information. Subsequently, we delve into the core modules of semantic vSLAM, including semantic extraction, object association, semantic loop closing, back-end optimization, and semantic mapping. Then, we present a performance comparison of semantic vSLAM systems under two different datasets, indoor and outdoor, respectively. Furthermore, we also provide a comparative analysis of widely used SLAM datasets to provide guidance for performance testing and validation. To further enrich the discussion, we identify unresolved challenges in semantic vSLAM, such as long-term semantic perception and association, open and unstructured environments. We propose future research directions, including balancing computational resources and quantifying system risk, large model-based navigation and mapping, and embodied AI SLAM. By providing key insights and forward-looking perspectives, this work aims to stimulate future research and improve the capabilities of semantic vSLAM in real-world applications.
Kaiqi Chen 0001, Junhao Xiao 0001, Qiyi Tong, Heng Zhang 0023, Ruyu Liu, Jianhua Zhang 0002, Arash Ajoudani, Shengyong Chen
IEEE Trans. Intell. Transp. Syst.1
2025 Covariance Propagation-Based Accurate Loop Detection for High Confusion Environment
Kaiqi Chen 0001, Ruyu Liu, Shengyong Chen, Arash Ajoudani, Jianhua Zhang 0002
IEEE Trans. Intell. Transp. Syst.2
2022 A Swift Gaze Estimate Method Based On The Corneal Image System
abstract
With the development of intelligent manufacturing, the demand of the incoming Human-Machine Interaction such as the augment reality rapidly increasing. However, the existing interaction modes in the augment reality, rely heavily on the hands or head movement. The inflexible modes is inefficient in the busy work flow. In this paper, we propose a gaze estimation work based on the Corneal Image System which can improve the efficiency of the interaction. Several prior works have proved, single Corneal Image contains the subject’s gaze information. However, the quality of Corneal Image is always impacted by the color and texture of the iris or the light from the surrounding, is hard to be applied directly. In order to improve the quality of the Corneal Image, people usually import additional devices into their work, such as infrared camera or eye tracker. These extra devices cause their gaze estimation works to become cumbersome and hard to be re-implemented commonly. Our gaze estimation work requires no additional device, can be seamlessly integrated into the AR domain with the help of the AprilTag mark. An AprilTag mark contained in an eye image, is distinct enough to be recognized, meanwhile, owns the hybrid pose relationship information between the eye, camera, and the focused AprilTag mark. The gaze can be inferred through the rigid body coordinate transformation naturally from this relationship. Many experiments have demonstrated that our approach is much easier to be re-implemented than the previous Corneal Image System based gaze computing works, at the same time, have the near performance to the state of the art.
Mengqi Du, Kaiqi Chen 0001, Jianhua Zhang 0002, Honghai Liu 0001
CSCWD2
2022 TXSLAM: A Monocular Semantic SLAM Tightly Coupled with Planar Text Features
abstract
We propose a new monocular semantic simultaneous localization and mapping (SLAM) system that tightly couples planar text features. The system treats text features as a plane with rich texture information and semantic information, and more accurate camera pose estimation can be obtained by tightly coupling the semantic plane. Unlike previous work, it pioneers the use of words contained in the text to represent the semantic information of the plane, which enables the use of simpler and more efficient data association algorithms to match geometric planes. We evaluate our method in public datasets, and the final experimental results prove that our proposed system improves the accuracy of camera pose estimation. Additionally, the system augments the sparse map with semantic plane information, enhancing the applicability of the system in robotics, unmanned driving, augmented reality (AR), and virtual reality (VR).
Qiyi Tong, Luzhen Ma, Kaiqi Chen 0001, Jianhua Zhang 0002
CSCWD3
2022 Robust and Accurate Multi-Agent SLAM with Efficient Communication for Smart Mobiles
abstract
In a long-term large-scenario application, the multi-agent collaborative SLAM is expected to improve the robustness and efficiency of executing tasks for mobile agents. In this paper, a multi-agent collaborative visual-inertial SLAM system is proposed based on a centralized client-server (CS) architecture, where the clients run on smart mobiles. In general, multi-agent collaborative SLAM relies on robust and precise experience sharing and efficient communication among agents. The experience sharing requires the place recognition with a high recall and accuracy, the precise estimation of transformation between looping frames, and the map fusion with globally consistency. To this end, we devise an enhanced geometric verification, a re-projection optimization based on the error-aware weighting strategy, and a strategy of flexible fusion to meet these requirements. In addition, the multi-agent collaborative SLAM needs to exchange abundant information, which requires the efficient communication. Therefore, we design a CS collaborative loop detection mechanism which is more robust to network transmission. We perform extensive experiments on the EuRoc dataset and in real environments. Experimental results show that the proposed system achieves better results than state-of-the-art methods. Furthermore, we demonstrate the stability of the proposed collaborative SLAM in real environments with a bandwidth of 7.55Mbps.
Kaiqi Chen 0001, Ruyu Liu, Yanhong Yang, Zhenhua Wang 0003, Jianhua Zhang 0002
ICRA2
2022 Robust Visual-Lidar Simultaneous Localization and Mapping System for UAV
abstract
Obtaining 3-D data by LIDAR from unmanned aerial vehicles (UAVs) is vital for the field of remote sensing; however, the highly dynamic movement of UAVs and narrow viewpoint of LIDAR pose a great challenge to the self-localization for UAVs based on solely LIDAR sensor. To this end, we propose a robust simultaneous localization and mapping (SLAM) system, which combines the image data obtained by vision sensor and point clouds obtained by LIDAR. In the front-end of the proposed system, the more stable line and plane features are extracted from point clouds through clustering. Then the relative pose between two consecutive frames is computed by the least squares iterative closest point algorithm. Afterward, a novel direct odometry algorithm is developed by combining the image frames and sparse point clouds, where the relative pose is used as a prior. In the back-end, the pose estimation is refined and the 3-D map with texture information is built at a lower frequency. Extensive experiments show that our method can achieve robust and highly precise localization and mapping for UAVs.
Kaiqi Chen 0001, Qinying Chen, Yanhong Yang, Jianhua Zhang 0002, Shengyong Chen
IEEE Geosci. Remote. Sens. Lett.2
2022 Accurate Object Association and Pose Updating for Semantic SLAM
abstract
Current pandemic has caused the medical system to operate under high load. To relieve it, robots with high autonomy can be used to effectively execute contactless operations in hospitals and reduce cross-infection between medical staff and patients. Although semantic Simultaneous Localization and Mapping (SLAM) technology can improve the autonomy of robots, semantic object association is still a problem that is worthy of being studied. The key to solving this problem is to correctly associate multiple object measurements of one object landmark by using semantic information, and to refine the pose of object landmark in real time. To this end, we propose a hierarchical object association strategy and a pose-refinement approach. The former one consists of two levels, i.e., a short-term object association and a global one. In the first level, we employ the multiple-object-tracking for short-term object association, through which the incorrect association among objects whose locations are close and appearances are similar can be avoided. Moreover, the short-term object association can provide more abundant object appearance and more robust estimation of object pose for the global object association in the second level. To refine the object pose in the map, we develop an approach to choose the optimal object pose from all object measurements associated with an object landmark. The proposed method is comprehensively evaluated on seven simulated hospital sequences, a real hospital environment and the KITTI dataset. Experimental results show that our method has an obviously improvement in terms of robustness and accuracy for the object association and the trajectory estimation in the semantic SLAM.
Kaiqi Chen 0001, Qinying Chen, Zhenhua Wang 0003, Jianhua Zhang 0002
IEEE Trans. Intell. Transp. Syst.1
2021 Collaborative Visual Inertial SLAM for Multiple Smart Phones
abstract
The efficiency and accuracy of mapping are crucial in a large scene and long-term AR applications. Multi-agent cooperative SLAM is the precondition of multi-user AR interaction. The cooperation of multiple smart phones has the potential to improve efficiency and robustness of task completion and can complete tasks that a single agent cannot do. However, it depends on robust communication, efficient location detection, robust mapping, and efficient information sharing among agents. We propose a multi-intelligence collaborative monocular visual-inertial SLAM deployed on multiple ios mobile devices with a centralized architecture. Each agent can independently explore the environment, run a visual-inertial odometry module online, and then send all the measurement information to a central server with higher computing resources. The server manages all the information received, detects overlapping areas, merges and optimizes the map, and shares information with the agents when needed. We have verified the performance of the system in public datasets and real environments. The accuracy of mapping and fusion of the proposed system is comparable to VINS-Mono which requires higher computing resources.
Ruyu Liu, Kaiqi Chen 0001, Jianhua Zhang 0002, Dongyan Guo
ICRA3
2021 Map Recovery and Fusion for Collaborative Augment Reality of Multiple Mobile Devices
abstract
The map recovery and fusion is a key issue in the application of large scale and long-term augmented reality (AR) scenarios. However, they are still not addressed well in an efficient and precise way, especially for complex industrial environments. In this article, we propose a map recovery and fusion strategy based on vision-inertial simultaneous localization and mapping. We first develop a heuristic strategy that can fast search and match map points among multiple maps, and can be used for efficient map fusion. For map recovery, we leverage the inertial sensors for short time motion estimation, and transform the previous lost map to the current map. Based on this strategy, a novel framework for collaborative AR is implemented and can parallelly run in multiple mobile devices in real time. Extensive experiments have been carried out on a public data set, and the results show that the proposed method can recovery and fuse multiple maps with high completeness and precision.
Jianhua Zhang 0002, Kaiqi Chen 0001, Zhiying Pan, Ruyu Liu, Thomas Yang 0001, Shengyong Chen
IEEE Trans. Ind. Informatics3