VLDB 2026 Research / reviewers in the wild / expert
Mingxing Wen
dblp:226/1703
· DBLP profile ↗
26ranked-venue papers
4as first author
17since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 22 · 4 first-author · 16 since 2021Systems, architecture and hardware · 12 · 1 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 3 first-author · 5 since 2021Databases, data management, data science and information retrieval · 3 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ChorusCVR: Chorus Supervision for Entire Space Post-Click Conversion Rate ModelingabstractPost-click conversion rate (CVR) estimation is a vital task in many recommender systems of revenue businesses, e.g., e-commerce and advertising. In a perspective of sample, a typical CVR positive sample usually goes through a funnel of exposure?click?conversion. For lack of post-event labels for un-clicked samples, CVR learning task commonly only utilizes clicked samples, rather than all exposed samples as for click-through rate (CTR) learning task. However, during online inference, CVR and CTR are estimated on the same assumed exposure space, which leads to a inconsistency of sample space between training and inference, i.e., sample selection bias (SSB). To alleviate SSB, previous wisdom proposes to design novel auxiliary tasks to enable the CVR learning on un-click training samples, such as CTCVR and counterfactual CVR, etc. Although alleviating SSB to some extent, none of them pay attention to the discrimination between ambiguous negative samples (un-clicked) and factual negative samples (clicked but un-converted) during modelling, which makes CVR model lacks robustness. To full this gap, we propose a novel ChorusCVR model to realize debiased CVR learning in entire-space. We propose a Negative sample Discrimination Module (NDM), which aims to provide robust soft labels with the ability to discriminate factual negative samples (clicked but un-converted) from ambiguous negative samples (un-clicked). Moreover, we propose a Soft Alignment Module (SAM) to supervise CVR learning with several alignment objectives using generated soft labels. Extensive offline experiments and online A/B testing at Kuaishou's e-commerce live service validates our ChorusCVR. Boyang Xia, Jiangxia Cao, Mingxing Wen, Zhaojie Liu, Liyin Hong, Kun Gai, Guorui Zhou |
WSDM | 6 |
| 2025 | QARM: Quantitative Alignment Multi-Modal Recommendation at KuaishouabstractIn recent years, with the significant evolution of multi-modal large models, many recommender researchers realized the potential of multi-modal information for user interest modeling. In industry, a wide-used modeling architecture is a cascading paradigm: (1) first pre-training a multi-modal model to provide omnipotent representations for downstream services; (2) The downstream recommendation model takes the multi-modal representation as additional input to fit real user-item behaviours. Although such paradigm achieves remarkable improvements, however, there still exist two problems that limit model performance: (1) Representation Unmatching: The pre-trained multi-modal model is always supervised by the classic NLP/CV tasks, while the recommendation models are supervised by real user-item interaction. As a result, the two fundamentally different tasks' goals were relatively separate, and there was a lack of consistent objective on their representations; (2) Representation Unlearning: The generated multi-modal representations are always stored in cache store and serve as extra fixed input of recommendation model, thus could not be updated by recommendation model gradient, further unfriendly for downstream training. Xinchen Luo, Jiangxia Cao, Jinkai Yu, Rui Huang 0009, Hezheng Lin, Yichen Zheng, Shiyao Wang 0001, Qigen Hu, Changqing Qiu, Xu Zhang 0065, Zhiheng Yan, Mingxing Wen, Zhaojie Liu, Guorui Zhou |
CIKM | 17 |
| 2025 | DSFormer-RTP: Dynamic-stream Transformers for Real-time Deterministic Trajectory PredictionabstractAs delivery robots are increasingly integrated into our daily lives, their ability to navigate through crowded spaces demands swift and accurate prediction of pedestrian trajectories, which is crucial for autonomous functionality. However, existing methods face challenges of unstable accuracy and inefficiency in real-world deployment. Trajectory prediction involves both temporal and social dimensions. Recent methods have achieved better results by modeling temporal and social dimensions simultaneously, preventing information loss compared to modeling them separately, which significantly increases computational costs, posing challenges for practical deployment.In this paper, we conceptualize the trajectory prediction task as a deterministic sequence-to-sequence model that produces one precise forecast, aligning with real-world needs while reducing complexity. To improve efficiency and reduce latency for real-time applications, we propose a novel dynamic-stream transformer architecture that categorizes layers into multi-stream and single-stream based on the number of dimensions involved in computation. The single-stream modules attend to all dimensions simultaneously, providing comprehensive information fusion but with higher computational complexity. The multi-stream modules focus on only one dimension, enabling parallel and batched computation, crucial for improving the model’s real-time performance. By combining them strategically, we achieve a balance between accuracy and speed. Extensive experiments on real datasets show that our dynamic-stream transformer architecture significantly reduces computational complexity, achieving a speed increase of 180% to 3180% compared to similar approaches, while also attaining performance close to the state-of-the-art (SOTA) for deterministic trajectory prediction. Mingxing Wen, Tianchen Deng, Danwei Wang |
IROS | 2 |
| 2024 | ACS-MM-Explore: Adaptive Circular Search Strategy for Multi-Modal Robot Exploration in Large-Scale Urban EnvironmentsabstractAutonomous exploration has become a crucial technology for mobile robots, and numerous broadly applicable algorithms have emerged. However, few exploration methods effectively utilize the features of specified types of areas to enhance the efficiency of autonomous exploration in a complex environment. In this paper, we propose ACS-MM-Explore, an adaptive-circular-search-based exploration framework for large-scale urban road environments, focusing on extracting and utilizing the boundaries of roads to enhance exploration efficiency. Our approach integrates a multi-modal traversabil-ity analysis module to distinguish between road and non-traversable areas on a 2D costmap. A novel mechanism for gen-erating exploration viewpoints is introduced, efficiently creating exploration viewpoints with a circular search process with an adaptive radius. An optimized viewpoint selection mechanism is included, taking into account the geographical and geomet-ric information of each viewpoint. The framework extends the move base and TEB local planner as a viewpoint-based navigation module. A comprehensive evaluation is concluded in a simulation environment, demonstrating the framework's effectiveness and robustness. Kaimin Mao, Mingxing Wen, Jun Zhang 0042, Guohao Peng, Zhenyu Wu 0001, Danwei Wang |
ICARCV | 3 |
| 2024 | OLIP-MIF: An Improved Method for Object Localization and Intention Prediction Based on Multimodal Information Fusionabstract3D object localization and intention prediction have become crucial components in autonomous system applications, such as self-driving car. However, there still faces a lot of challenges, especially for complex and dynamic scenarios where a single modality information is insufficient to effectively and precisely localize the position and analyze the intention of objects. An improved method based on multimodal information fusion has been proposed via leveraging the advantages of 2D image segmentation and 3D geometrical characteristics of LiDAR point cloud. Extensive comparative experiments have been conducted and the results demonstrate that the proposed method significantly enhances both localization and prediction accuracy, comparing with the method where 2D bounding box of object instead of segmentation information is used to be fused with point cloud. Mingxing Wen, Hongmiaoyi Zhang, Jinwei Huang, Shuomin Huang, Yunyao Lyv, Yisheng Guan, Danwei Wang |
ICARCV | 1 |
| 2024 | Real-Time GNSS Spoofing Detection for Autonomous Vehicles: An Attention-Based Autoencoder ApproachabstractWith the rapid evolution of autonomous vehicles (AVs), ensuring reliable navigation has become paramount, especially against threats like Global Navigation Satellite Systems (GNSS) spoofing. This paper presents an attention-based Autoencoder approach for real-time GNSS spoofing detection in AVs. The proposed method leverages data from multiple sensors, including IMU, GNSS, and LiDAR, fully utilizing the redundancy and correlations among them. By integrating a multi-head attention mechanism into the Autoencoder, the model can thoroughly capture and analyze the complex relationships within sensor data, enhancing its capability to promptly and accurately identify spoofing attacks. Field experiments demonstrate the effectiveness of the proposed method in achieving a high detection rate and short detection time, highlighting its potential for practical deployment in AV applications. Mingxing Wen, Yuanzhe Wang |
ICARCV | 4 |
| 2024 | LB-R2R-Calib: Accurate and Robust Extrinsic Calibration of Multiple Long Baseline 4D Imaging Radars for V2XabstractAs a new sensor, 4D radar (x, y, z, velocity) has great potential for V2X, due to its 3D point cloud, direct doppler velocity output, long distance ranging, low-cost, and more importantly, robust perception in all weathers. However, the extrinsic calibration of multiple long baseline 4D radars is rarely researched in V2X, which is the key to fuse multi-radars. The main reasons are three-folds: (1) New sensor. Thus, it is not surprising that little related work can be found. (2) Long baseline and large viewpoint-difference. Current works are mainly focused on unmanned vehicles, which is short baseline and small viewpoint-difference. (3) Sparse, noisy, and very cluttered 4D radar point cloud. Thus, it is challenging to rapidly and accurately locate the target and extract the feature. In this paper, LB-R2R-Calib (Long Baseline Radar to Radar extrinsic Calibration) is proposed to address these problems. The novelties are: (1) A new target is introduced: an eight-quadrant corner reflector enclosed by a foam sphere. The benefit is the target center is a viewpoint-invariant feature. Thus, it is ideal for large viewpoint-difference calibration. (2) A new feature extraction algorithm is proposed to rapidly locate the target and extract the target center from a very cluttered point cloud, as we observed some important characteristics of 4D radar. Experiments with two 4D radars in real environments with four configurations demonstrate our method is highly accurate and robust. Jun Zhang 0042, Fangwei Zhang, Zhenyu Wu 0001, Guohao Peng, Yiyao Liu, Qiyang Lyu, Mingxing Wen, Danwei Wang |
ICRA | 8 |
| 2023 | CAHIR: Co-Attentive Hierarchical Image Representations for Visual Place RecognitionabstractRobust visual place recognition (VPR) against significant appearance changes is crucial for the life-long operation of mobile robots. Focusing on this task, we propose a Co-Attentive Hierarchical Image Representations (CAHIR) framework for VPR, which unifies attention-sharing global and local descriptor generation into one encoding pipeline. The hierarchical descriptors are applied to a coarse-to-fine VPR system with global retrieval and local geometric verification. To explore high-quality local matches between task-relevant visual elements, a cross-attention mutual enhancement layer is introduced to strengthen the information interaction between the local descriptors. Through the proposed selective matching distillation, the mutual enhancement layer can learn from state-of-the-art local matchers in a distillation manner. After weighted cross-matching of the enhanced local descriptors, geometric verification is applied to evaluate the spatial consistency of the compared image pair. Experiments show CAHIR outperforms the existing global and local representations for VPR in terms of performance and efficiency. Quantitatively, it achieves state-of-the-art results on three city-scale benchmark datasets. Qualitatively, CAHIR proves to attach great importance to task-relevant visual elements and excels at finding local correspondences that are discriminative to the VPR task. Guohao Peng, Heshan Li, Jun Zhang 0042, Mingxing Wen, Singh Rahul, Danwei Wang |
ICRA | 5 |
| 2023 | 4DRadarSLAM: A 4D Imaging Radar SLAM System for Large-scale Environments based on Pose Graph OptimizationabstractLiDAR-based SLAM may easily fail in adverse weathers (e.g., rain, snow, smoke, fog), while mmWave Radar remains unaffected. However, current researches are primarily focused on 2D$(x,y)$or 3D ($x, y$, doppler) Radar and 3D LiDAR, while limited work can be found for 4D Radar ($x, y, z$, doppler). As a new entrant to the market with unique characteristics, 4D Radar outputs 3D point cloud with added elevation information, rather than 2D point cloud; compared with 3D LiDAR, 4D Radar has noisier and sparser point cloud, making it more challenging to extract geometric features (edge and plane). In this paper, we propose a full system for 4D Radar SLAM consisting of three modules: 1) Front-end module performs scan-to-scan matching to calculate the odometry based on GICP, considering the probability distribution of each point; 2) Loop detection utilizes multiple rule-based loop pre-filtering steps, followed by an intensity scan context step to identify loop candidates, and odometry check to reject false loop; 3) Back-end builds a pose graph using front-end odometry, loop closure, and optional GPS data. Optimal pose is achieved through$\mathrm{g}2\mathrm{o}$. We conducted real experiments on two platforms and five datasets (ranging from 240m to 4.8km) and will make the code open-source to promote further research at: https://github.com/zhuge2333/4DRadarSLAM Jun Zhang 0042, Huayang Zhuge, Zhenyu Wu 0001, Guohao Peng, Mingxing Wen, Yiyao Liu, Danwei Wang |
ICRA | 5 |
| 2023 | LB-L2L-Calib 2.0: A Novel Online Extrinsic Calibration Method for Multiple Long Baseline 3D LiDARs Using ObjectsabstractIn V2X (Vehicle-to-Everything), one important work is to extrinsically calibrate multiple 3D LiDARs, which are mounted with a long baseline and large viewpoint-difference at the road-side. Current solutions either require a specific target being set up (e.g., a sphere), or require specific features existing in the environment (e.g., mutually orthogonal planes). However, it is time-consuming, sometimes even inconvenient, to set up specific targets, e.g., at busy intersections and highways. Furthermore, specific features do not always exist in the traffic scenario. Thus, the current solutions are not feasible. To address this problem, a novel extrinsic calibration method is proposed in this paper, namely LB-L2L-Calib 2.0. It is the 2.0 version of our previous work. The novelties are: 1) We propose to use the easily accessible objects on the road as features for calibration (i.e., the vehicles). Thus, it is not necessary to set up any specific targets and we do not need to worry whether specific features exist or not. The key point is we observed that the 3D bounding box centers of the vehicles are viewpoint-invariant from different viewpoints, which makes them ideal features for long baseline and large viewpoint-difference calibration. 2) To establish correct correspondence between the bounding box centers detected from different LiDARs, we propose an exhaustive searching strategy. It can robustly output correct correspondence. Extensive experiments are performed in three scenarios (simulation: intersection, real: carpark and highway), with two types of LiDAR (Velodyne and Livox), demonstrating that LB-L2L-Calib 2.0 is robust, effective, and accurate. Jun Zhang 0042, Qiao Yan, Mingxing Wen, Qiyang Lyu, Guohao Peng, Zhenyu Wu 0001, Danwei Wang |
IROS | 3 |
| 2023 | L2V2T2Calib: Automatic and Unified Extrinsic Calibration Toolbox for Different 3D LiDAR, Visual Camera and Thermal CameraabstractExtrinsic calibration between LiDAR-Camera and LiDAR-LiDAR has been researched extensively, because it is the foundation for sensor fusion. Meanwhile, many projects are open-sourced and significantly promote related research. However, limited solutions can unify the calibration between repetitive scanning and non-repetitive scanning 3D LiDAR, sparse and dense 3D LiDAR, visual and thermal camera. Currently, to achieve that, we normally need to use different targets and extract different features for different sensor combinations. Sometimes, human intervention is required to locate the target. It is inconvenient and time-consuming. In this paper, L2V2T2Calib is introduced and open-sourced as a trial to unify the calibration. 1). A four-circular-holes board is adopted for all sensors. The four circle centers can be detected by all the sensors, thus are ideal common features. Previous works also use this target, but the algorithms don’t consider non-repetitive scanning LiDARs, thus cannot be directly applied. 2). To unify the process, an important step is to automatically and robustly detect the target from different types of LiDARs. However, this does not receive enough attention. We propose a method based on template matching. It is simple, but effective and general to different depth sensors. 3). We provide two types of output, minimizing 2D re-projection error (Min2D) and minimizing 3D matching error (Min3D), for different users. And their performance is compared. Extensive experiments conducted in both simulation and real environment demonstrate L2V2T2Calib is accurate, robust, more importantly, unified. The code will be open-sourced to promote related research at: https://github.com/Clothooo/lvt2calib Jun Zhang 0042, Yiyao Liu, Mingxing Wen, Yufeng Yue, Danwei Wang |
IV | 3 |
| 2023 | Integrated Localization and Planning for Cruise Control of UGV Platoons in Infrastructure-Free EnvironmentsabstractThis paper investigates the cruise control problem of unmanned ground vehicle (UGV) platoons from the implementation perspective. Unlike most existing works related to platoon cruise control which rely on positioning infrastructures such as lane markings, roadside units, and global navigation satellite systems (GNSS), this paper explores a new problem: platoon cruise control in environments without positioning infrastructures. The introduction of this constraint disables most existing cruise control approaches. To address this problem, an integrated localization and planning framework is proposed, which is composed of three modular algorithms. Firstly, to localize multiple vehicles in a common coordinate system, a collaborative localization algorithm is developed through matching local perceptions of different vehicles. Secondly, to maintain the desired platoon configuration, the historical trajectory of the preceding vehicle is reconstructed, based on which the target state is planned for the following vehicle. Finally, a virtual controller based algorithm is designed to generate feasible trajectories for the following vehicle in real time. The proposed framework has two salient features. Firstly, it does not depend on positioning infrastructures and does not introduce additional positioning sensors, such as GNSS/INS modules, ultra-wideband (UWB) devices, magnetic meters and so on, as long as each vehicle is equipped with a perception sensor (Lidar, radar or camera), which however is essential equipment for nowaday autonomous systems. Secondly, the proposed framework does not depend on direct observations between vehicles to achieve relative localization, making it applicable in non-line-of-sight (non-LOS) situations. Real-world experiments have been conducted to validate the effectiveness, robustness and practicality of the proposed framework. Yuanzhe Wang, Mingxing Wen, Yufeng Yue, Danwei Wang |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2022 | C-TM: Topo-metric Mapping and Localization based on Place Categorization and Place Recognition for a Delivery Robot on FootpathabstractIn this work, C-TM is presented: a method to build a topo-metric map for delivery robot navigation in largescale city environments. This system automatically generates a compact map by only saving expensive LIDAR information at key locations. These locations form the nodes of a topological map. Nodes are identified using a Place-Categorization (PC) neural network which output the place category from RGB cameras. Inside nodes, we generate and save high quality LIDAR submaps. Global localization within the map is done with a Visual-Place-Recognition (VPR) neural network. The topo-metric map can be used for navigation on footpath. We deploy C-TM on a four-wheeled autonomous delivery robot and test the effectiveness in two environments, both day and night. Timothy Chia, Jun Zhang 0042, Heshan Li, Guohao Peng, Mingxing Wen, Dawei Kee, P. G. C. N. Senarathne |
ICARCV | 5 |
| 2022 | Vision Based Sidewalk Navigation for Last-mile Delivery RobotabstractNavigating delivery robot along the sidewalk safely and robustly in a campus environment is extremely challenging due to the narrow motion space, appearance changes and unstable GPS localization signal under canopies of trees, etc. To that end, we have completed a systematic implementation for delivery robot sidewalk navigation, where a robust vision based navigation algorithm has been proposed. And it consists of three main modules: sidewalk segmentation, costmap generation and motion planning. More Specifically, the first module is to find the drivable area of the surrounding environment, where an image-based segmentation neural network has been developed to extract where the robot can traverse. Since it only takes as input immediate and local sensory data, thus releasing the high dependence on a prior map. Then, an inverse perspective mapping follows to generate a bird-eye-view of the drivable area and constructs the local occupancy grid map intuitively. Next, two different motion planners, control-based primitives (Dynamic Window Approach) and state-based primitives (state lattice planner), have been adopted to generate a trajectory candidate for navigating the robot along the sidewalk. Both simulation and real-world sidewalk navigation experiments have been conducted to test and evaluate their performance. The results show that our algorithm can precisely extract the sidewalk area for traversing, and the state-based primitive planner demonstrates superior performance in terms of trajectory length and time cost, achieving 14.3% and 18.7% improvement compared with control-based primitive planner. Mingxing Wen, Jun Zhang 0042, Tairan Chen, Guohao Peng, Timothy Chia, Yingchong Ma |
ICARCV | 1 |
| 2022 | A Robust Sidewalk Navigation Method for Mobile Robots Based on Sparse Semantic Point CloudabstractLast-mile delivery robots are usually required to navigate on the sidewalk through a fixed route. The current solutions heavily rely on the image-based perception and GPS localization to successfully complete delivery tasks. However, it is prone to fail and become unreliable when the robot runs in challenging conditions, such as operating in different illuminations, or under canopies of trees or buildings. To address these issues, this paper proposes a novel robust sidewalk navigation method for the last-mile delivery robots with an affordable sparse LiDAR, which consists of two main modules: Semantic Point Cloud Network (SegPCn) and Reactive Nav-igation Network (RNn), as shown in Fig. 1. More specifically, SegPCn takes the raw 3D point cloud as input and predicts the point-wise segmentation labels, presenting a robust perception capability even in the night. Then, the semantic point clouds are fed to RNn to generate an angular velocity to navigate the robot along the sidewalk, where the localization of the robot is not required. Moreover, an autolabeling mechanism is developed to reduce the labor involved in data preparation as well. And the LSTM neural network is explored to effectively leverage the historical context and derive correct decisions. Extensive experiments have been carried out to verify the efficacy of this method, and the results show that this method enables the robot to navigate on the sidewalk robustly during day and night. We open source the code and the data set on https://github.com/lukewenMX/Robust-Navigation-Method. Mingxing Wen, Yunxiang Dai, Tairan Chen, Jun Zhang 0042, Danwei Wang |
IROS | 1 |
| 2021 | MSTSL: Multi-Sensor Based Two-Step Localization in Geometrically Symmetric EnvironmentsabstractSymmetric environment is one of the most intractable and challenging scenarios for mobile robots to accomplish global localization tasks, due to the highly similar geometrical structures and insufficient distinctive features. Existing localization solutions in such scenarios either depend on pre-deployed infrastructures which are expensive, inflexible, and hard to maintain; or rely on single sensor-based methods whose initialization module is incapable to provide enough unique information. Thus, this paper proposes a novel Multi-Sensor based Two-Step Localization framework named MSTSL, which addresses the problem of mobile robot global localization in geometrically symmetric environments by utilizing the measured magnetic field, 2-D LiDAR, and wheel odometry information. The proposed system mainly consists of two steps: 1) Magnetic Field-based Initialization, and 2) LiDAR-based Localization. Based on the pre-built magnetic field database, multiple initial hypotheses poses can firstly be determined by the proposed two-stage initialization algorithm. Then, utilizing the obtained multiple initial hypotheses, the robot can be localized more accurately by LiDAR-based localization. Extensive experiments demonstrate the practical utility and accuracy of the proposed system over the alternative approaches in real-world scenarios. Zhenyu Wu 0001, Yufeng Yue, Mingxing Wen, Jun Zhang 0042, Guohao Peng, Danwei Wang |
ICRA | 3 |
| 2021 | Tightly-Coupled Perception and Navigation of Heterogeneous Land-Air Robots in Complex ScenariosabstractIn unstructured and unknown environments, heterogeneous robots must be able to perceive the environment, coordinate with each other and complete tasks collaboratively with onboard sensors. In this paper, a tightly-coupled perception and navigation framework is proposed for heterogeneous land-air robots, which forms a closed loop of perception-navigation for heterogeneous robots. The key novelty of this work is the proposing of a unified framework to formulate the cooperative mapping and navigation problem, as well as the derivation of high-level coordination strategy and low-level goal-oriented navigation within a fully integrated approach. To provide a comprehensive understanding of the environment, a flexible probabilistic map fusion algorithm is applied to merge local maps generated by hybrid robots. The proposed UAV-UGV hybrid system is validated in challenging experiments, proving its robustness and effectiveness in practical tasks. Yufeng Yue, Mingxing Wen, Yosmar Putra, Meiling Wang 0002, Danwei Wang |
ICRA | 2 |
| 2020 | HILPS: Human-in-Loop Policy Search for Mobile Robot NavigationabstractReinforcement learning has obtained increasing attention in mobile robot mapless navigation in recent years. However, there are still some obvious challenges including the sample efficiency, safety due to dilemma of exploration and exploitation. These problems are addressed in this paper by proposing the Human-in-Loop Policy Search (HILPS) framework, where learning from demonstration, learning from human intervention and Near Optimal Policy strategies are integrated together. Firstly, the former two make sure that expert experience grant mobile robot a more informative and correct decision for accomplishing the task and also maintaining the safety of the mobile robot due to the priority of human control. Then the Near Optimal Policy (NOP) provides a way to selectively store the similar experience with respect to the preexisting human demonstration, in which case the sample efficiency can be improved by eliminating exclusively exploratory behaviors. To verify the performance of the algorithm, the mobile robot navigation experiments are extensively conducted in simulation and real world. Results show that HILPS can improve sample efficiency and safety in comparison to state-of-art reinforcement learning. Mingxing Wen, Yufeng Yue, Zhenyu Wu 0001, Ehsan Mihankhah, Danwei Wang |
ICARCV | 1 |
| 2020 | Multi-Robot Collaborative Reasoning for Unique Person Recognition in Complex EnvironmentsabstractThe discovery of unique or suspicious people is essential for active surveillance of security or patrol robots, and multi-robot collaboration and dynamic reasoning can further enhance their adaptability in large-scale environments. This paper proposes a hierarchical probabilistic reasoning framework for a multi-robot system to actively identify the unique person with distinct motion patterns in large-scale and dynamic environments. Linear and angular velocities are considered typical motion patterns, which are extracted by using heterogeneous sensors to detect and track people. First, single robot reasoning is performed, each robot judges the uniqueness of people by comparing their motion patterns based on local observations. Meanwhile, multi-robot reasoning is also performed, by fusing the perceptual information from each individual robot to form a global observation and then make another judgment based on it. Finally, each robot can decide which result should be adopted by comparing the beliefs of local and global judgments. Experimental results show that the method is feasible in various environments. Chule Yang, Yufeng Yue, Mingxing Wen, Yuanzhe Wang |
ICARCV | 3 |
| 2020 | Day and Night Collaborative Dynamic Mapping in Unstructured Environment Based on Multimodal SensorsabstractEnabling long-term operation during day and night for collaborative robots requires a comprehensive understanding of the unstructured environment. Besides, in the dynamic environment, robots must be able to recognize dynamic objects and collaboratively build a global map. This paper proposes a novel approach for dynamic collaborative mapping based on multimodal environmental perception. For each mission, robots first apply heterogeneous sensor fusion model to detect humans and separate them to acquire static observations. Then, the collaborative mapping is performed to estimate the relative position between robots and local 3D maps are integrated into a globally consistent 3D map. The experiment is conducted in the day and night rainforest with moving people. The results show the accuracy, robustness, and versatility in 3D map fusion missions. Yufeng Yue, Chule Yang, Jun Zhang 0042, Mingxing Wen, Zhenyu Wu 0001, Danwei Wang |
ICRA | 4 |
| 2020 | A Hierarchical Framework for Collaborative Probabilistic Semantic MappingabstractPerforming collaborative semantic mapping is a critical challenge for cooperative robots to maintain a comprehensive contextual understanding of the surroundings. Most of the existing work either focus on single robot semantic mapping or collaborative geometry mapping. In this paper, a novel hierarchical collaborative probabilistic semantic mapping framework is proposed, where the problem is formulated in a distributed setting. The key novelty of this work is the mathematical modeling of the overall collaborative semantic mapping problem and the derivation of its probability decomposition. In the single robot level, the semantic point cloud is obtained based on heterogeneous sensor fusion model and is used to generate local semantic maps. Since the voxel correspondence is unknown in collaborative robots level, an Expectation-Maximization approach is proposed to estimate the hidden data association, where Bayesian rule is applied to perform semantic and occupancy probability update. The experimental results show the high quality global semantic map, demonstrating the accuracy and utility of 3D semantic map fusion algorithm in real missions. Yufeng Yue, Chule Yang, Jun Zhang 0042, Mingxing Wen, Yuanzhe Wang, Danwei Wang |
ICRA | 6 |
| 2020 | Infrastructure-Free Global Localization in Repetitive Environments: An OverviewabstractRepetitive environment is a challenging scenario for mobile robot global localization due to its highly similar structures and lack of distinctive features. Existing solutions in such environments rely heavily on pre-installed infrastructures, which are neither flexible nor cost-effective. Besides, few of the previous research have been focused on the implementation of infrastructure-free localization approaches in repetitive scenarios. Thus, this paper serves as a survey to investigate the problem of infrastructure-free mobile robot global localization with low-cost and efficient sensors in repetitive environments. Three of the most popular infrastructure-free localization methods, namely LiDAR-based localization (LBL), vision-based localization (VBL), and magnetic field-based localization (MFL), are analyzed and evaluated. Extensive global localization experiments are conducted in real-world repetitive scenarios and the results demonstrate that VBL methods perform slightly better than LBL and MFL methods. The overall evaluations indicate that infrastructure-free global localization in repetitive environment is still a challenging problem which deserves more research efforts to develop new solutions. Zhenyu Wu 0001, Jun Zhang 0042, Yufeng Yue, Mingxing Wen, Zichen Jiang, Danwei Wang |
IECON | 4 |
| 2020 | Collaborative Semantic Perception and Relative Localization Based on Map MatchingabstractIn order to enable a team of robots to operate successfully, retrieving accurate relative transformation between robots is the fundamental requirement. So far, most research on relative localization mainly focus on geometry features such as points, lines and planes. To address this problem, collaborative semantic map matching is proposed to perform semantic perception and relative localization. This paper performs semantic perception, probabilistic data association and nonlinear optimization within an integrated framework. Since the voxel correspondence between partial maps is a hidden variable, a probabilistic semantic data association algorithm is proposed based on Expectation-Maximization. Instead of specifying hard geometry data association, semantic and geometry association are jointly updated and estimated. The experimental verification on Semantic KITTI benchmarks demonstrate the improved robustness and accuracy. Yufeng Yue, Mingxing Wen, Zhenyu Wu 0001, Danwei Wang |
IROS | 3 |
| 2019 | Probabilistic Reasoning for Unique Role Recognition Based on the Fusion of Semantic-Interaction and Spatio-Temporal FeaturesabstractThis paper deals with the problem of recognizing the unique role in dynamic environments. Different from social roles, the unique role refers to those who are unusual in their carrying items or movements in the scene. In this paper, we propose a hierarchical probabilistic reasoning method that relates spatial relationships between interested objects and humans with their temporal changes to recognize the unique individual. Two observation models, Object Existence Model (OEM) and Human Action Model (HAM), are established to support role inference by analyzing the corresponding semantic-interaction features and spatio-temporal features. Then, OEM and HAM results of each person are compared with the overall distribution in the scene, respectively. Finally, we can determine the role through the fusion of two observation models. Experiments are conducted in both indoor and outdoor environments concerning different settings, degrees of clutter, and occlusions. The results show that the proposed method can adapt to a variety of scenarios and outperforms other methods on accuracy and robustness, moreover, exhibiting stable performance even in complex scenes. Chule Yang, Yufeng Yue, Jun Zhang 0042, Mingxing Wen, Danwei Wang |
IEEE Trans. Multim. | 4 |
| 2018 | Probabilistic Fusion Framework for Collaborative Robots 3D MappingabstractFusion of local 3D maps generated by individual robots to a globally consistent 3D map is one of the fundamental challenges in multi-robot mapping missions. In this paper, we propose a probabilistic mathematical formulation to address the integrated map fusion problem. More specifically, the problem of estimating fused map posterior can be factorized into a product of relative transformation posterior and the global map posterior, which enables us to solve map matching and map merging problems efficiently. In addition, a distributed communication strategy is employed to share map information among robots. The proposed approach is evaluated in indoor and mixed environments, which shows its utility in 3D map fusion for multi-robot mapping missions. Yufeng Yue, P. G. C. N. Senarathne, Chule Yang, Jun Zhang 0042, Mingxing Wen, Danwei Wang |
FUSION | 5 |
| 2018 | A Two-step Method for Extrinsic Calibration between a Sparse 3D LiDAR and a Thermal CameraabstractTo obtain the 6 DOF extrinsic parameters (rotation and translation matrix) between a 3D ranging sensor and a thermal camera, previous methods require a high-resolution 3D ranging sensor to reliably detect features. Although sparse 3D LiDARs are widely used on autonomous robots, to the best of our knowledge, the extrinsic calibration between a sparse 3D LiDAR (particularly Velodyne VLP-16) and a thermal camera has not been considered in the literature. In this paper, we present a two-step method to address the problem, where a monocular visual camera is used to assist the process. The proposed method decomposes the problem into two steps: extrinsic calibration between a sparse 3D LiDAR and a visual camera; extrinsic calibration between a visual camera and a thermal camera. Experiments are conducted to demonstrate the effectiveness of the proposed two-step method. Jun Zhang 0042, Prarinya Siritanawan, Yufeng Yue, Chule Yang, Mingxing Wen, Danwei Wang |
ICARCV | 5 |