VLDB 2026 Research / reviewers in the wild / expert
Yifan Duan
dblp:231/3980
· DBLP profile ↗
21ranked-venue papers
5as first author
20since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 15 · 5 first-author · 15 since 2021Systems, architecture and hardware · 9 · 2 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 5 since 2021Computer networks · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | NuwaDynamics+: A Causality-Aware Generative Framework for Spatio-Temporal Representation LearningabstractSpatio-temporal (ST) prediction is crucial in earth sciences, including meteorological forecasting and urban computing, to name just a few. Access to ample high-quality data, combined with deep models adept at inference, is essential for attaining significant outcomes. Yet, data scarcity and the substantial costs of sensor deployment result in notable data imbalances. Overly specialized models that lack causal linkages further undermine the generalizability of inference techniques. To address these challenges, we first introduce a causal framework for ST predictions, named $\mathtt{NuwaDynamics}$NuwaDynamicsNuwaDynamics, aimed at pinpointing causal regions in data and providing models with the capability for causal reasoning in a dual-phase process. Initially, we employ upstream self-supervision to identify causally significant patches, equipping the model with generalizable insights and performing targeted interventions on non-essential patches to approximate potential testing distributions. This stage is known as the discovery phase. Progressing from discovery, we apply the insights to downstream tasks tailored to specific ST goals, enhancing the model's recognition of a wider potential data distribution and augmenting its causal perceptual abilities (referred to as the Update phase). Additionally, we address environmental controllability and high computational complexity by implementing channel multiplication and conditional generation methods. This process, termed $\mathtt{NuwaDynamics+}$NuwaDynamics+NuwaDynamics+, can further be interpreted as the front-door adjustment technique in the causality domain. Through comprehensive experiments across ten real-world or simulated ST benchmarks, we demonstrate that integrating the $\mathtt{NuwaDynamics+}$NuwaDynamics+NuwaDynamics+ concept substantially improves various model performance. $\mathtt{NuwaDynamics+}$NuwaDynamics+NuwaDynamics+ concept also significantly enhances the versatility across various dynamic ST tasks, such as extreme weather forecasting and long-temporal-step super-resolution predictions. Kun Wang 0056, Yifan Duan, Hao Wu 0083, Jian Zhao 0006, Kai Wang 0036, Zhengyang Zhou, Yuxuan Liang 0002, Xu Wang 0029, Yang Wang 0015, Yu Zheng 0004, Xuelong Li 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2025 | RaCFormer: Towards High-Quality 3D Object Detection via Query-based Radar-Camera FusionabstractWe propose Radar-Camera fusion transformer (RaC-Former) to boost the accuracy of 3D object detection by the following insight. The Radar-Camera fusion in outdoor 3D scene perception is capped by the image-to-BEV transformation-if the depth of pixels is not accurately estimated, the naive combination of BEV features actually integrates unaligned visual content. To avoid this problem, we propose a query-based framework that enables adaptive sampling of instance-relevant features from both the bird’s-eye view (BEV) and the original image view. Furthermore, we enhance system performance by two key designs: optimizing query initialization and strengthening the representational capacity of BEV. For the former, we introduce an adaptive circular distribution in polar coordinates to refine the initialization of object queries, allowing for a distance-based adjustment of query density. For the latter, we initially incorporate a radar-guided depth head to refine the transformation from image view to BEV. Subsequently, we focus on leveraging the Doppler effect of radar and introduce an implicit dynamic catcher to capture the temporal elements within the BEV. Extensive experiments on nuScenes and View-of-Delft (VoD) datasets validate the merits of our design. Remarkably, our method achieves superior results of 64.9% mAP and 70.2% NDS on nuScenes. RaCFormer also secures the state-of-the-art performance on the VoD dataset. Code is available at https://github.com/cxmomo/RaCFormer. Xiaomeng Chu, Jiajun Deng, Guoliang You, Yifan Duan, Houqiang Li, Yanyong Zhang |
CVPR | 4 |
| 2025 | ORPP: Self-Optimizing Role-playing Prompts to Enhance Language Model CapabilitiesabstractHigh-quality prompts are crucial for eliciting outstanding performance from large language models (LLMs) on complex tasks.Existing research has explored model-driven strategies for prompt optimization.However, these methods often suffer from high computational overhead or require strong optimization capabilities from the model itself, which limits their broad applicability.To address these challenges, we propose ORPP, a framework that enhances model performance by optimizing and generating roleplaying prompts.The core idea of ORPP is to confine the prompt search space to role-playing scenarios, thereby fully activating the model's intrinsic capabilities through carefully crafted, high-quality role-playing prompts.Specifically, ORPP first performs iterative optimization on a small subset of training samples to generate high-quality role-playing prompts.Then, leveraging the model's few-shot learning capability, it transfers the optimization experience to efficiently generate suitable prompts for the remaining samples.Our experimental results show that ORPP not only matches but in most cases surpasses existing mainstream prompt optimization methods in terms of performance.Notably, ORPP suggests great "plug-and-play" capability.In most cases, it can be integrated with various other prompt methods and further enhance their effectiveness. Yifan Duan, Yihong Tang, Kehai Chen, Liqiang Nie, Min Zhang 0005 |
EMNLP | 1 |
| 2025 | CELLmap: Enhancing LiDAR SLAM Through Elastic and Lightweight Spherical Map RepresentationabstractSLAM is a fundamental capability of unmanned systems, with LiDAR-based SLAM gaining widespread adoption due to its high precision. Current SLAM systems can achieve centimeter-level accuracy within a short period. However, there are still several challenges when dealing with largescale mapping tasks including significant storage requirements and difficulty of reusing the constructed maps. To address this, we first design an elastic and lightweight map representation called CELLmap, composed of several CELLS, each representing the local map at the corresponding location. Then, we design a general backend including CELL-based bidirectional registration module and loop closure detection module to improve global map consistency. Our experiments have demonstrated that CELLmap can represent the precise geometric structure of large-scale maps of KITTI dataset using only about 60 MB. Additionally, our general backend achieves up to a 26.88% improvement over various LiDAR odometry methods. Yifan Duan, Yao Li 0016, Guoliang You, Xiaomeng Chu, Jianmin Ji, Yanyong Zhang |
ICRA | 1 |
| 2025 | OG-Gaussian: Occupancy Based Street Gaussians for Autonomous DrivingabstractAccurate and realistic 3D scene reconstruction enables the lifelike creation of autonomous driving simulation environments. With advancements in 3D Gaussian Splatting (3DGS), previous studies have applied it to reconstruct complex dynamic driving scenes. These methods typically require expensive LiDAR sensors and pre-annotated datasets of dynamic objects. To address these challenges, we propose OG-Gaussian, a novel approach that replaces LiDAR point clouds with Occupancy Grids (OGs) generated from surround-view camera images using Occupancy Prediction Network (ONet). Our method leverages the semantic information in OGs to separate dynamic vehicles from static street background, converting these grids into two distinct sets of initial point clouds for reconstructing both static and dynamic objects. Additionally, we estimate the trajectories and poses of dynamic objects through a learning-based approach, eliminating the need for complex manual annotations. Experiments on Waymo Open dataset demonstrate that OG-Gaussian is on par with the current state-of-the-art in terms of reconstruction quality and rendering speed, achieving an average PSNR of 35.13 and a rendering speed of 143 FPS, while significantly reducing computational costs and economic overhead. Yedong Shen, Yifan Duan, Yilong Wu, Jianmin Ji, Yanyong Zhang, Huiqing Jin |
ICRA | 3 |
| 2025 | MT-PCR: Leveraging Modality Transformation for Large-Scale Point Cloud Registration with Limited OverlapabstractLarge-scale scene point cloud registration with limited overlap is a challenging task due to computational load and constrained data acquisition. To tackle these issues, we propose a point cloud registration method, MT-PCR, based on Modality Transformation. MT-PCR leverages a Bird's Eye View (BEV) capturing the maximal overlap information to improve the accuracy and utilizes images to provide complementary spatial features. Specifically, MT-PCR converts 3D point clouds to BEV images and estimates correspondence by 2D image keypoints extraction and matching. Subsequently, the 2D correspondence estimates are then transformed back to 3D point clouds using inverse mapping. We have applied MT-PCR to Terrestrial Laser Scanning (TLS) and Aerial Laser Scanning (ALS) point cloud registration on the GrAco dataset, involving 8 low-overlap, square-kilometer scale registration scenarios. Experiments and comparisons with commonly used methods demonstrate that MT-PCR can achieve superior accuracy and robustness in large-scale scenes with limited overlap. Yilong Wu, Yifan Duan, Yedong Shen, Jianmin Ji, Yanyong Zhang |
ICRA | 2 |
| 2025 | CalibWorkflow: A General MLLM-Guided Workflow for Centimeter-Level Cross-Sensor CalibrationabstractExtrinsic calibration is a fundamental step in sensor fusion systems. However, existing methods often lack generalization capabilities when facing diverse hardware configurations, sensor poses, and environmental conditions, hindering their large-scale deployment. To address this limitation, we propose a general extrinsic calibration method, CalibWorkflow. Our core innovation lies in positioning multimodal large language models (MLLMs) as ''visual guides'' for the calibration process, leveraging their powerful vision-language understanding capabilities to guide parameter search and refinement. This reliance on visual scene understanding, rather than specific geometric features or sensor characteristics, enables the method to generalize effectively across diverse hardware and environmental conditions. Specifically, CalibWorkflow employs a three-stage calibration pipeline: initial parameter search, coarse optimization, and fine optimization. First, it utilizes the MLLM to assess the visual consistency between the projected point cloud and the image, rapidly determining an initial range for the extrinsic parameters. Next, the MLLM serves as a differential evaluator, giving simple ''better'' or ''worse'' feedback on parameter changes to guide the search through the parameter space. Finally, the method refines the calibration by matching edge features and performing non-linear optimization. Extensive experiments are conducted across six diverse scenarios and four heterogeneous sensor combinations. CalibWorkflow achieves state-of-the-art sub-degree and centimeter-level accuracy on four datasets and demonstrates highly competitive performance on others. These results thoroughly validate the generalization and robustness when facing various scenarios. Codes will be available. Wuyang Zhang, Guoliang You, Xiaomeng Chu, Wenhao Yu 0010, Yifan Duan, Yanyong Zhang |
ACM Multimedia | 6 |
| 2025 | GS-Share: Enabling High-fidelity Map Sharing with Incremental Gaussian SplattingabstractAbstract Constructing and sharing 3D maps is essential for many applications, including autonomous driving and augmented reality. Recently, 3D Gaussian splatting has emerged as a promising approach for accurate 3D reconstruction. However, a practical map‐sharing system that features high‐fidelity, continuous updates, and network efficiency remains elusive. To address these challenges, we introduce GS‐Share, a photorealistic map‐sharing system with a compact representation. The core of GS‐Share includes anchor‐based global map construction, virtual‐image‐based map enhancement, and incremental map update. We evaluate GS‐Share against state‐of‐the‐art methods, demonstrating that our system achieves higher fidelity, particularly for extrapolated views, with improvements of 11%, 22%, and 74% in PSNR, LPIPS, and Depth L1, respectively. Furthermore, GS‐Share is significantly more compact, reducing map transmission overhead by 36%. Hanqi Zhu, Yifan Duan, Yanyong Zhang |
Comput. Graph. Forum | 3 |
| 2025 | A novel anomaly detection and classification algorithm for application in tuyere images of blast furnace
Yifan Duan, Yanqin Sun |
Eng. Appl. Artif. Intell. | 1 |
| 2024 | NuwaDynamics: Discovering and Updating in Causal Spatio-Temporal ModelingabstractSpatio-temporal (ST) prediction plays a pivotal role in earth sciences, such as meteorological prediction, urban computing. Adequate high-quality data, coupled with deep models capable of inference, are both indispensable and prerequisite for achieving meaningful results. However, the sparsity of data and the high costs associated with deploying sensors lead to significant data imbalances. Models that are overly tailored and lack causal relationships further compromise the generalizabilities of inference methods. Towards this end, we first establish a causal concept for ST predictions, named NuwaDynamics, which targets to identify causal regions in data and endow model with causal reasoning ability in a two-stage process. Concretely, we initially leverage upstream self-supervision to discern causal important patches, imbuing the model with generalized information and conducting informed interventions on complementary trivial patches to extrapolate potential test distributions. This phase is referred to as the discovery step. Advancing beyond discovery step, we transfer the data to downstream tasks for targeted ST objectives, aiding the model in recognizing a broader potential distribution and fostering its causal perceptual capabilities (refer as Update step). Our concept aligns seamlessly with the contemporary backdoor adjustment mechanism in causality theory. Extensive experiments on six real-world ST benchmarks showcase that models can gain outcomes upon the integration of the NuwaDynamics concept. NuwaDynamics also can significantly benefit a wide range of changeable ST tasks like extreme weather and long temporal step super-resolution predictions. Kun Wang 0056, Hao Wu 0083, Yifan Duan, Guibin Zhang, Kai Wang 0036, Xiaojiang Peng, Yu Zheng 0004, Yuxuan Liang 0002, Yang Wang 0015 |
ICLR | 3 |
| 2024 | OCC-VO: Dense Mapping via 3D Occupancy-Based Visual Odometry for Autonomous DrivingabstractVisual Odometry (VO) plays a pivotal role in autonomous systems, with a principal challenge being the lack of depth information in camera images. This paper introduces OCC-VO, a novel framework that capitalizes on recent advances in deep learning to transform 2D camera images into 3D semantic occupancy, thereby circumventing the traditional need for concurrent estimation of ego poses and landmark locations. Within this framework, we utilize the TPV-Former to convert surround view cameras’ images into 3D semantic occupancy. Addressing the challenges presented by this transformation, we have specifically tailored a pose estimation and mapping algorithm that incorporates Semantic Label Filter, Dynamic Object Filter, and finally, utilizes Voxel PFilter for maintaining a consistent global semantic map. Evaluations on the Occ3D-nuScenes not only showcase a 20.6% improvement in Success Ratio and a 29.6% enhancement in trajectory accuracy against ORB-SLAM3, but also emphasize our ability to construct a comprehensive map. Our implementation is open-sourced and available at: https://github.com/USTCLH/OCC-VO. Yifan Duan, Jianmin Ji, Yanyong Zhang |
ICRA | 2 |
| 2024 | LDP: A Local Diffusion Planner for Efficient Robot Navigation and Collision AvoidanceabstractThe conditional diffusion model has been demonstrated as an efficient tool for learning robot policies, owing to its advancement to accurately model the conditional distribution of policies. The intricate nature of real-world scenarios, characterized by dynamic obstacles and maze-like structures, underscores the complexity of robot local navigation decision-making as a conditional distribution problem. Nevertheless, leveraging the diffusion model for robot local navigation is not trivial and encounters several under-explored challenges: (1) Data Urgency The complex conditional distribution in local navigation needs training data to include diverse policy in diverse real-world scenarios; (2) Myopic Observation Due to the diversity of the perception scenarios, diffusion decisions based on the local perspective of robots may prove suboptimal for completing the entire task, as they often lack foresight. In certain scenarios requiring detours, the robot may become trapped. To address these issues, our approach begins with an exploration of a diverse data generation mechanism that encompasses multiple agents exhibiting distinct preferences through target selection informed by integrated global-local insights. Then, based on this diverse training data, a diffusion agent is obtained, capable of excellent collision avoidance in diverse scenarios. Subsequently, we augment our Local Diffusion Planner, also known as LDP by incorporating global observations in a lightweight manner. This enhancement broadens the observational scope of LDP, effectively mitigating the risk of becoming ensnared in local optima and promoting more robust navigational decisions. Our experimental results demonstrated that the LDP outperforms other baseline algorithms in navigation performance, exhibiting enhanced robustness across diverse scenarios with different policy preferences and superior generalization capabilities for unseen scenarios. Moreover, we highlighted the competitive advantage of the LDP within real-world settings. Wenhao Yu 0010, Jie Peng 0002, Junrui Zhang 0012, Yifan Duan, Jianmin Ji, Yanyong Zhang |
IROS | 5 |
| 2024 | CRPlace: Camera-Radar Fusion with BEV Representation for Place RecognitionabstractThe integration of complementary characteristics from camera and radar data has emerged as an effective approach in 3D object detection. However, such fusion-based methods remain unexplored for place recognition, an equally important task for autonomous systems. Given that place recognition relies on the similarity between a query scene and the corresponding candidate scene, the stationary background of a scene is expected to play a crucial role in the task. As such, current well-designed camera-radar fusion methods for 3D object detection can hardly take effect in place recognition because they mainly focus on dynamic foreground objects. In this paper, a background-attentive camera-radar fusion-based method, named CRPlace, is proposed to generate background-attentive global descriptors from multi-view images and radar point clouds for accurate place recognition. To extract stationary background features effectively, we design an adaptive module that generates the background-attentive mask by utilizing the camera BEV feature and radar dynamic points. With the guidance of a background mask, we devise a bidirectional cross-attention-based spatial fusion strategy to facilitate comprehensive spatial interaction between the background information of the camera BEV feature and the radar BEV feature. As the first camera-radar fusion-based place recognition network, CRPlace has been evaluated thoroughly on the nuScenes dataset. The results show that our algorithm outperforms a variety of baseline methods across a comprehensive set of metrics (recall@1 reaches 91.2%). Shaowei Fu, Yifan Duan, Yao Li 0016, Chengzhen Meng, Jianmin Ji, Yanyong Zhang |
IROS | 2 |
| 2024 | MM-Gaussian: 3D Gaussian-based Multi-modal Fusion for Localization and Reconstruction in Unbounded ScenesabstractLocalization and mapping are critical tasks for various applications such as autonomous vehicles and robotics. The challenges posed by outdoor environments present particular complexities due to their unbounded characteristics. In this work, we present MM-Gaussian, a LiDAR-camera multimodal fusion system for localization and mapping in unbounded scenes. Our approach is inspired by the recently developed 3D Gaussians, which demonstrate remarkable capabilities in achieving high rendering quality and fast rendering speed. Specifically, our system fully utilizes the geometric structure information provided by solid-state LiDAR to address the problem of inaccurate depth encountered when relying solely on visual solutions in unbounded, outdoor scenarios. Additionally, we utilize 3D Gaussian point clouds, with the assistance of pixel-level gradient descent, to fully exploit the color information in photos, thereby achieving realistic rendering effects. To further bolster the robustness of our system, we designed a relocalization module, which assists in returning to the correct trajectory in the event of a localization failure. Experiments conducted in multiple scenarios demonstrate the effectiveness of our method. Yifan Duan, Yu Sheng, Jianmin Ji, Yanyong Zhang |
IROS | 2 |
| 2024 | RayFormer: Improving Query-Based Multi-Camera 3D Object Detection via Ray-Centric StrategiesabstractThe recent advances in query-based multi-camera 3D object detection are featured by initializing object queries in the 3D space, and then sampling features from perspective-view images to perform multi-round query refinement. In such a framework, query points near the same camera ray are likely to sample similar features from very close pixels, resulting in ambiguous query features and degraded detection accuracy. To this end, we introduce RayFormer, a camera-ray-inspired query-based 3D object detector that aligns the initialization and feature extraction of object queries with the optical characteristics of cameras. Specifically, RayFormer transforms perspective-view image features into bird's eye view (BEV) via the lift-splat-shoot method and segments the BEV map to sectors based on the camera rays. Object queries are uniformly and sparsely initialized along each camera ray, facilitating the projection of different queries onto different areas in the image to extract distinct features. Besides, we leverage the instance information of images to supplement the uniformly initialized object queries by further involving additional queries along the ray from 2D object detection boxes. To extract unique object-level features that cater to distinct queries, we design a ray sampling method that suitably organizes the distribution of feature sampling points on both images and bird's eye view. Extensive experiments are conducted on the nuScenes dataset to validate our proposed ray-inspired model design. The proposed RayFormer achieves 55.5% mAP and 63.3% NDS, respectively. Xiaomeng Chu, Jiajun Deng, Guoliang You, Yifan Duan, Yao Li 0016, Yanyong Zhang |
ACM Multimedia | 4 |
| 2024 | Map++: Towards User-Participatory Visual SLAM Systems with Efficient Map Expansion and SharingabstractConstructing precise 3D maps is crucial for the development of future map-based systems such as self-driving and navigation. However, generating these maps in complex environments, such as multi-level parking garages or shopping malls, remains a formidable challenge. In this paper, we introduce a participatory sensing approach that delegates map-building tasks to map users, thereby enabling cost-effective and continuous data collection. The proposed method harnesses the collective efforts of users, facilitating the expansion and ongoing update of the maps as the environment evolves. Hanqi Zhu, Yifan Duan, Wuyang Zhang, Longfei Shangguan, Yu Zhang 0086, Jianmin Ji, Yanyong Zhang |
MobiCom | 3 |
| 2024 | Causal Deciphering and Inpainting in Spatio-Temporal Dynamics via Diffusion ModelabstractSpatio-temporal (ST) prediction has garnered a De facto attention in earth sciences, such as meteorological prediction, human mobility perception. However, the scarcity of data coupled with the high expenses involved in sensor deployment results in notable data imbalances. Furthermore, models that are excessively customized and devoid of causal connections further undermine the generalizability and interpretability. To this end, we establish a causal framework for ST predictions, termed CaPaint, which targets to identify causal regions in data and endow model with causal reasoning ability in a two-stage process. Going beyond this process, we utilize the back-door adjustment to specifically address the sub-regions identified as non-causal in the upstream phase. Specifically, we employ a novel image inpainting technique. By using a fine-tuned unconditional Diffusion Probabilistic Model (DDPM) as the generative prior, we in-fill the masks defined as environmental parts, offering the possibility of reliable extrapolation for potential data distributions. CaPaint overcomes the high complexity dilemma of optimal ST causal discovery models by reducing the data generation complexity from exponential to quasi-linear levels. Extensive experiments conducted on five real-world ST benchmarks demonstrate that integrating the CaPaint concept allows models to achieve improvements ranging from 4.3% to 77.3%. Moreover, compared to traditional mainstream ST augmenters, CaPaint underscores the potential of diffusion models in ST enhancement, offering a novel paradigm for this field. Our project is available at https://anonymous.4open.science/r/12345-DFCC. Yifan Duan, Jian Zhao 0006, pengcheng, Junyuan Mao, Hao Wu 0098, Jingyu Xu 0002, Shilong Wang 0002, Caoyuan Ma, Kai Wang 0036, Kun Wang 0056, Xuelong Li 0001 |
NeurIPS | 1 |
| 2023 | P3O: Transferring Visual Representations for Reinforcement Learning via PromptingabstractIt is important for deep reinforcement learning (DRL) algorithms to transfer their learned policies to new environments that have different visual inputs. In this paper, we introduce Prompt based Proximal Policy Optimization (P3O), a three-stage DRL algorithm that transfers visual representations from a target to a source environment by applying prompting. The process of P3O consists of three stages: pre-training, prompting, and predicting. In particular, we specify a prompt-transformer for representation conversion and propose a two-step training process to train the prompt-transformer for the target environment, while the rest of the DRL pipeline remains unchanged. We implement P3O and evaluate it on the OpenAI CarRacing video game. The experimental results show that P3O outperforms the state-of-the-art visual transferring schemes. In particular, P3O allows the learned policies to perform well in environments with different visual inputs, which is much more effective than retraining the policies in these environments. Guoliang You, Xiaomeng Chu, Yifan Duan, Jie Peng 0002, Jianmin Ji, Yu Zhang 0086, Yanyong Zhang |
ICME | 3 |
| 2022 | PFilter: Building Persistent Maps through Feature Filtering for Fast and Accurate LiDAR-based SLAMabstractSimultaneous localization and mapping (SLAM) based on laser sensors has been widely adopted by mobile robots and autonomous vehicles. These SLAM systems are required to support accurate localization with limited computational resources. In particular, point cloud registration, i.e., the process of matching and aligning multiple LiDAR scans collected at multiple locations in a global coordinate framework, has been deemed as the bottleneck step in SLAM. In this paper, we propose a feature filtering algorithm, PFilter, that can filter out invalid features and can thus greatly alleviate this bottleneck. Meanwhile, the overall registration accuracy is also improved due to the carefully curated feature points. We integrate PFilter into the well-established scan-to-map LiDAR odometry framework, F-LOAM, and evaluate its performance on the KITTI dataset. The experimental results show that PFilter can remove about 48.4% of the points in the local feature map and reduce feature points in scan by 19.3% on average, which save 20.9% processing time per frame. In the mean time, we improve the accuracy by 9.4%. Yifan Duan, Jie Peng 0002, Yu Zhang 0086, Jianmin Ji, Yanyong Zhang |
IROS | 1 |
| 2021 | Towards an Online RRT-based Path Planning Algorithm for Ackermann-steering VehiclesabstractIt is challenging to develop an online path planning algorithm for Ackermann-steering vehicles to find collision-free and kinematically-feasible paths, that is efficient for dense environments, adaptable to various environments, and suitable for environments with narrow passages. In this paper, we propose a kinematically constrained RRT-based path planning algorithm integrating with a trajectory parameter space (TP-space) with three novel improvements to meet the above requirements. In specific, we introduce a new way to choose candidate nodes to expand the tree for an RRT-based algorithm, which can significantly increase the success rate of the expansion and improve the efficiency of the algorithm. We also introduce a procedure to incrementally adjust the step size for the expansion, which enables the algorithm to automatically adapt to various environments. At last, we integrate rapidly-exploring random vines (RRV) with a TP-space to handle kinematic constraints and improve the performance of the algorithm to expand the tree through a narrow passage. We also prove that the algorithm is probabilistic complete and asymptotically near-optimal. An ablation study shows that all three improvements can notably improve the performance of the RRT-based path planning algorithm. We also evaluate the algorithm in various environments. The experimental results show that our algorithm achieves competitive performance compared with the state-of-the-art. The source code is available at https://github.com/PengJieb/fastbkrrt. Jie Peng 0002, Yu'an Chen, Yifan Duan, Yu Zhang 0086, Jianmin Ji, Yanyong Zhang |
ICRA | 3 |
| 2019 | Cross-Layer Design for Mission-Critical IoT in Mobile Edge Computing SystemsabstractIn this paper, we establish a cross-layer framework for optimizing user association, packet offloading rates, and bandwidth allocation for mission-critical Internet-of-Things (MC-IoT) services with short packets in mobile edge computing (MEC) systems, where enhanced mobile broadband (eMBB) services with long packets are considered as background services. To reduce communication delay, the fifth generation new radio is adopted in radio access networks. To avoid long queueing delay for short packets from MC-IoT, processor-sharing (PS) servers are deployed at MEC systems, where the service rate of the server is equally allocated to all the packets in the buffer. We derive the distribution of latency experienced by short packets in closed form, and minimize the overall packet loss probability subject to the end-to-end delay requirement. To solve the nonconvex optimization problem, we propose an algorithm that converges to a near optimal solution when the throughput of eMBB services is much higher than MC-IoT services, and extend it into more general scenarios. Furthermore, we derive the optimal solutions in two asymptotic cases: communication or computing is the bottleneck of reliability. The simulation and numerical results validate our analysis and show that the PS server outperforms first-come-first-serve servers. Changyang She, Yifan Duan, Guodong Zhao 0001, Tony Q. S. Quek, Yonghui Li 0001, Branka Vucetic |
IEEE Internet Things J. | 2 |