EDBT 2026 Demo / reviewers in the wild / expert
Guangyao Zhai
dblp:243/2753
· DBLP profile ↗
22ranked-venue papers
7as first author
17since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 19 · 6 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 2 first-author · 9 since 2021Systems, architecture and hardware · 5 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | HALT: Hierarchical Attention Learning for Visual TrackingabstractSiamese tracking algorithms have gained widespread recognition due to their exceptional efficiency and scalability. However, they exhibit suboptimal performance when localizing arbitrary targets under various disturbances, particularly in complex environments involving challenges such as illumination variation, deformation, and background clutter. Therefore, this paper proposes a Hierarchical Attention Learning network (HAL) to enhance tracking performance. Drawing inspiration from the hybrid attention mechanism, HAL designs a Feature Mining Attention module (FMA), a Global Feature Attention module (GFA), and a hierarchical network structure. Concretely, FMA employs parallel branches to fully extract channel and spatial features, enabling preliminary enhancement of target features and establishing global associations. Since lower-level and higher-level features emphasize positional and semantic information, respectively, GFA comprehensively integrates multi-level features to obtain more accurate tracking predictions. In particular, to improve the model's representation capacity, a hierarchical network structure is developed to deepen the HAL network and strengthen the feature dependencies between the template and the search region. Finally, based on HAL, we propose a Hierarchical Attention Learning Tracker (HALT) for visual tracking in complex environments, which is capable of learning rich hierarchical features. Extensive experiments demonstrate that, compared to state-of-the-art trackers, our HALT achieves outstanding tracking performance across multiple benchmarks while maintaining a real-time speed of 48.5 fps. Fengwei Gu, Ao Li 0002, Chen Chen 0086, Chengtao Cai, Renjie Qiao, Guangyao Zhai, Zhaojie Ju |
IEEE Trans Autom. Sci. Eng. | 8 |
| 2025 | MMGDreamer: Mixed-Modality Graph for Geometry-Controllable 3D Indoor Scene GenerationabstractControllable 3D scene generation has extensive applications in virtual reality and interior design, where the generated scenes should exhibit high levels of realism and controllability in terms of geometry. Scene graphs provide a suitable data representation that facilitates these applications. However, current graph-based methods for scene generation are constrained to text-based inputs and exhibit insufficient adaptability to flexible user inputs, hindering the ability to precisely control object geometry. To address this issue, we propose MMGDreamer, a dual-branch diffusion model for scene generation that incorporates a novel Mixed-Modality Graph, visual enhancement module, and relation predictor. The mixed-modality graph allows object nodes to integrate textual and visual modalities, with optional relationships between nodes. It enhances adaptability to flexible user inputs and enables meticulous control over the geometry of objects in the generated scenes. The visual enhancement module enriches the visual fidelity of text-only nodes by constructing visual representations using text embeddings. Furthermore, our relation predictor leverages node representations to infer absent relationships between nodes, resulting in more coherent scene layouts. Extensive experimental results demonstrate that MMGDreamer exhibits superior control of object geometry, achieving state-of-the-art scene generation performance. Zhifei Yang 0004, Keyang Lu, Jiaxing Qi, Hanqi Jiang, Ruifei Ma, Shenglin Yin, Yifan Xu 0028, Mingzhe Xing, Jieyi Long, Xiangde Liu, Guangyao Zhai |
AAAI | 13 |
| 2025 | KGN-Pro: Keypoint-Based Grasp Prediction through Probabilistic 2D-3D Correspondence LearningabstractHigh-level robotic manipulation tasks demand flexible 6-DoF grasp estimation to serve as a basic function. Previous approaches either directly generate grasps from point-cloud data, suffering from challenges with small objects and sensor noise, or infer 3D information from RGB images, which introduces expensive annotation requirements and discretization issues. Recent methods mitigate some challenges by retaining a 2D representation to estimate grasp keypoints and applying Perspective-n-Point (PnP) algorithms to compute 6-DoF poses. However, these methods are limited by their non-differentiable nature and reliance solely on 2D supervision, which hinders the full exploitation of rich 3D information. In this work, we present KGN-Pro, a novel grasping network that preserves the efficiency and fine-grained object grasping of previous KGNs while integrating direct 3D optimization through probabilistic PnP layers. KGN-Pro encodes paired RGB-D images to generate Keypoint Map, and further outputs a 2D confidence map to weight keypoint contributions during re-projection error minimization. By modeling the weighted sum of squared re-projection errors probabilistically, the network effectively transmits 3D supervision to its 2D keypoint predictions, enabling end-to-end learning. Experiments on both simulated and real-world platforms demonstrate that KGN-Pro outperforms existing methods in terms of grasp cover rate and success rate. Project website: https://waitderek.github.io/kgnpro. Bingran Chen, Baorun Li, Jian Yang 0003, Yong Liu 0007, Guangyao Zhai |
IROS | 5 |
| 2025 | Video Perception Models for 3D Scene SynthesisabstractAutomating the expert-dependent and labor-intensive task of 3D scene synthesis would significantly benefit fields such as architectural design, robotics simulation, and virtual reality. Recent approaches to 3D scene synthesis often rely on the commonsense reasoning of large language models (LLMs) or strong visual priors from image generation models. However, current LLMs exhibit limited 3D spatial reasoning, undermining the realism and global coherence of synthesized scenes, while image-generation-based methods often constrain viewpoint control and introduce multi-view inconsistencies. In this work, we present Video Perception models for 3D Scene synthesis (VIPScene), a novel framework that exploits the encoded commonsense knowledge of the 3D physical world in video generation models to ensure coherent scene layouts and consistent object placements across views. VIPScene accepts both text and image prompts and seamlessly integrates video generation, feedforward 3D reconstruction, and open-vocabulary perception models to semantically and geometrically analyze each object in a scene. This enables flexible scene synthesis with high realism and structural consistency. For a more sufficient evaluation on coherence and plausibility, we further introduce First-Person View Score (FPVScore), utilizing a continuous first-person perspective to capitalize on the reasoning ability of multimodal large language models. Extensive experiments show that VIPScene significantly outperforms existing methods and generalizes well across diverse scenarios. Rui Huang 0012, Guangyao Zhai, Zuria Bauer, Marc Pollefeys, Federico Tombari, Leonidas J. Guibas, Gao Huang 0001, Francis Engelmann |
NeurIPS | 2 |
| 2024 | SecondPose: SE(3)-Consistent Dual-Stream Feature Fusion for Category-Level Pose EstimationabstractCategory-level object pose estimation, aiming to predict the 6D pose and 3D size of objects from known categories, typically struggles with large intra-class shape variation. Existing works utilizing mean shapes often fall short of cap-turing this variation. To address this issue, we present Sec-ondPose, a novel approach integrating object-specific ge-ometric features with semantic category priors from DI-NOv2. Leveraging the advantage of DINOv2 in providing SE(3)-consistent semantic features, we hierarchically extract two types of SE(3)-invariant geometric features to further encapsulate local-to-global object-specific information. These geometric features are then point-aligned with DINOv2 features to establish a consistent object represen-tation under SE(3) transformations, facilitating the map-ping from camera space to the pre-defined canonical space, thus further enhancing pose estimation. Extensive exper-iments on NOCS-REAL275 demonstrate that SecondPose achieves a 12.4% leap forward over the state-of-the-art. Moreover, on a more complex dataset HouseCat6D which provides photometrically challenging objects, SecondPose still surpasses other competitors by a large margin. Yamei Chen, Yan Di, Guangyao Zhai, Fabian Manhardt, Chenyangguang Zhang, Ruida Zhang, Federico Tombari, Nassir Navab, Benjamin Busam |
CVPR | 3 |
| 2024 | ShapeMatcher: Self-Supervised Joint Shape Canonicalization, Segmentation, Retrieval and DeformationabstractIn this paper, we present ShapeMatcher, a unified self-supervised learning framework for joint shape canonicalization, segmentation, retrieval and deformation. Given a partially-observed object in an arbitrary pose, we first canonicalize the object by extracting point-wise affine-invariant features, disentangling inherent structure of the object with its pose and size. These learned features are then leveraged to predict semantically consistent part segmentation and corresponding part centers. Next, our lightweight retrieval module aggregates the features within each part as its retrieval token and compare all the tokens with source shapes from a pre-established database to identify the most geometrically similar shape. Finally, we deform the retrieved shape in the deformation module to tightly fit the input object by harnessing part center guided neural cage deformation. The key insight of ShapeMaker is the simultaneous training of the four highly-associated processes: canonicalization, segmentation, retrieval, and deformation, leveraging cross-task consistency losses for mutual supervision. Extensive experiments on synthetic datasets PartNet, ComplementMe, and real-world dataset Scan2CAD demonstrate that ShapeMatcher surpasses competitors by a large margin. Code is released at https://github.com/Det1999/ShapeMaker. Yan Di, Chenyangguang Zhang, Chaowei Wang, Ruida Zhang, Guangyao Zhai, Xiangyang Ji, Shan Gao 0003 |
CVPR | 5 |
| 2024 | HouseCat6D - A Large-Scale Multi-Modal Category Level 6D Object Perception Dataset with Household Objects in Realistic ScenariosabstractEstimating 6D object poses is a major challenge in 3D computer vision. Building on successful instance-level approaches, research is shifting towards category-level pose estimation for practical applications. Current category-level datasets, however, fall short in annotation quality and pose variety. Addressing this, we introduce HouseCat6D, a new category-level 6D pose dataset. It features 1) multi-modality with Polarimetric RGB and Depth (RGBD+P), 2) encompasses 194 diverse objects across 10 household cat-egories, including two photometrically challenging ones, and 3) provides high-quality pose annotations with an error range of only 1.35 mm to 1.74 mm. The dataset also includes 4) 41 large-scale scenes with comprehensive view-point and occlusion coverage,5) a checkerboard-free en-vironment, and 6) dense 6D parallel-jaw robotic grasp annotations. Additionally, we present benchmark results for leading category-level pose estimation networks. Patrick Ruhkamp, Guangyao Zhai, Hannah Schieber, Giulia Rizzoli, Pengyuan Wang 0002, Hongcheng Zhao, Lorenzo Garattoni, Daniel Roth 0001, Sven Meier, Nassir Navab, Benjamin Busam |
CVPR | 4 |
| 2024 | GeoGaussian: Geometry-Aware Gaussian Splatting for Scene Rendering
Yanyan Li 0001, Chenyu Lyu, Yan Di, Guangyao Zhai, Gim Hee Lee, Federico Tombari |
ECCV (35) | 4 |
| 2024 | EchoScene: Indoor Scene Generation via Information Echo Over Scene Graph Diffusion
Guangyao Zhai, Evin Pinar Örnek, Dave Zhenyu Chen, Ruotong Liao, Yan Di, Nassir Navab, Federico Tombari, Benjamin Busam |
ECCV (21) | 1 |
| 2024 | SG-Bot: Object Rearrangement via Coarse-to-Fine Robotic Imagination on Scene GraphsabstractObject rearrangement is pivotal in robotic-environment interactions, representing a significant capability in embodied AI. In this paper, we present SG-Bot, a novel rearrangement framework that utilizes a coarse-to-fine scheme with a scene graph as the scene representation. Unlike previous methods that rely on either known goal priors or zero-shot large models, SG-Bot exemplifies lightweight, real-time, and user-controllable characteristics, seamlessly blending the consideration of commonsense knowledge with automatic generation capabilities. SG-Bot employs a three-fold procedure– observation, imagination, and execution–to adeptly address the task. Initially, objects are discerned and extracted from a cluttered scene during the observation. These objects are first coarsely organized and depicted within a scene graph, guided by either commonsense or user-defined criteria. Then, this scene graph subsequently informs a generative model, which forms a fine-grained goal scene considering the shape information from the initial scene and object semantics. Finally, for execution, the initial and envisioned goal scenes are matched to formulate robotic action policies. Experimental results demonstrate that SG-Bot outperforms competitors by a large margin. Guangyao Zhai, Xiaoni Cai, Dianye Huang, Yan Di, Fabian Manhardt, Federico Tombari, Nassir Navab, Benjamin Busam |
ICRA | 1 |
| 2024 | DMA-Net: A dual branch encoder and multi-scale cross attention fusion network for skin lesion segmentationabstractAbstract Automatic segmentation of skin lesion is an important step in computer‐aided diagnosis. However, due to the significant variations in the size and shape of the lesion areas, as well as the low contrast with normal skin tissue, the boundaries are not clearly distinguishable, leading to a high possibility of incorrect segmentation. Therefore, this task is highly challenging. To overcome these difficulties, this paper proposes a medical image segmentation architecture named dual branch encoder and multi‐scale cross attention fusion network, which includes a dual‐branch encoder based on convolutional neural network and an improved channel‐enhanced Mamba to comprehensively extract local and global information from dermoscopy images. Additionally, to enhance the feature interaction and fusion of local and global information, a multi‐scale cross attention fusion module is adopted to cross‐merge features in different directions and at different scales, maximizing the advantages of the dual‐branch encoder and achieving precise segmentation of skin lesions. Extensive experiments are conducted on three public skin lesion datasets: ISIC‐2018, ISIC‐2017, and ISIC‐2016, to verify the effectiveness and superiority of the proposed method. The dice similarity coefficient scores on the three datasets reached 81.77%, 81.68% and 85.60%, respectively, surpassing most state‐of‐the‐art methods. Guangyao Zhai, Guanglei Wang 0002, Qinghua Shang, Yan Li 0053, Hongrui Wang 0002 |
IET Image Process. | 1 |
| 2024 | Enabling Versatility and Dexterity of the Dual-Arm Manipulators: A General Framework Toward Universal Cooperative ManipulationabstractGrasping and manipulating various kinds of objects cooperatively is the core skill of a dual-arm robot when deployed as an autonomous agent in a human-centered environment. This requires fully exploiting the robot's versatility and dexterity. In this work, we propose a general framework for dual-arm manipulators that contains two correlative modules. The learning-based dexterity-reachability-aware perception module deals with vision-based bimanual grasping. It employs an end-to-end evaluation network and probabilistic modeling of the robot's reachability to deliver feasible and dexterity-optimum grasp pairs for unseen objects. The optimization-based versatility-oriented control module addresses the online cooperative manipulation control by using a hierarchical quadratic programming formulation. Self-collision avoidance and dual-arm manipulability ellipsoid tracking with high reliability and fidelity are simultaneously achieved based on a learned lightweight distance proxy function and a speed-level tracking technique on Riemannian manifold. Intrinsic system safety is guaranteed, and a novel interface for skill transfer is enabled. A long-horizon rearrangement experiment, a bimanual turnover manipulation, and multiple comparative performance evaluation verify the effectiveness of the proposed framework. Zhehua Zhou, Yang Yang 0031, Guangyao Zhai, Marion Leibold, Fenglei Ni, Zhengyou Zhang, Martin Buss, Yu Zheng 0001 |
IEEE Trans. Robotics | 5 |
| 2023 | On the Importance of Accurate Geometry Data for Dense 3D Vision TasksabstractLearning-based methods to solve dense 3D vision problems typically train on 3D sensor data. The respectively used principle of measuring distances provides advantages and drawbacks. These are typically not compared nor discussed in the literature due to a lack of multi-modal datasets. Texture-less regions are problematic for structure from motion and stereo, reflective material poses issues for active sensing, and distances for translucent objects are intricate to measure with existing hardware. Training on inaccurate or corrupt data induces model bias and hampers generalisation capabilities. These effects remain unnoticed if the sensor measurement is considered as ground truth during the evaluation. This paper investigates the effect of sensor errors for the dense 3D vision tasks of depth estimation and reconstruction. We rigorously show the significant impact of sensor characteristics on the learned predictions and notice generalisation issues arising from various technologies in everyday household environments. For evaluation, we introduce a carefully designed dataset11dataset available at https://github.com/Junggy/HAMMER-dataset comprising measurements from commodity sensors, namely D-ToF, I-ToF, passive/active stereo, and monocular RGB+P. Our study quantifies the considerable sensor noise impact and paves the way to improved dense vision estimates and targeted data fusion. Patrick Ruhkamp, Guangyao Zhai, Nikolas Brasch, Yannick Verdie, Jifei Song, Yiren Zhou, Anil Armagan, Slobodan Ilic, Ales Leonardis, Nassir Navab, Benjamin Busam |
CVPR | 3 |
| 2023 | IPCC-TP: Utilizing Incremental Pearson Correlation Coefficient for Joint Multi-Agent Trajectory PredictionabstractReliable multi-agent trajectory prediction is crucial for the safe planning and control of autonomous systems. Compared with single-agent cases, the major challenge in simultaneously processing multiple agents lies in modeling complex social interactions caused by various driving intentions and road conditions. Previous methods typically leverage graph-based message propagation or attention mechanism to encapsulate such interactions in the format of marginal probabilistic distributions. However, it is inherently sub-optimal. In this paper, we propose IPCC-TP, a novel relevance-aware module based on Incremental Pearson Correlation Coefficient to improve multi-agent interaction modeling. IPCC-TP learns pairwise joint Gaussian Distributions through the tightly-coupled estimation of the means and covariances according to interactive incremental movements. Our module can be conveniently embedded into existing multi-agent prediction methods to extend original motion distribution decoders. Extensive experiments on nuScenes and Argoverse 2 datasets demonstrate that IPCC-TP improves the performance of baselines by a large margin. Dekai Zhu, Guangyao Zhai, Yan Di, Fabian Manhardt, Hendrik Berkemeyer, Nassir Navab, Federico Tombari, Benjamin Busam |
CVPR | 2 |
| 2023 | MonoGraspNet: 6-DoF Grasping with a Single RGB Imageabstract6-DoF robotic grasping is a long-lasting but un-solved problem. Recent methods utilize strong 3D networks to extract geometric grasping representations from depth sensors, demonstrating superior accuracy on common objects but performing unsatisfactorily on photometrically challenging objects, e.g., objects in transparent or reflective materials. The bottleneck lies in that the surface of these objects can not reflect accurate depth due to the absorption or refraction of light. In this paper, in contrast to exploiting the inaccurate depth data, we propose the first RGB-only 6-DoF grasping pipeline called MonoGraspNet that utilizes stable 2D features to simultaneously handle arbitrary object grasping and overcome the problems induced by photometrically challenging objects. MonoGraspNet leverages a keypoint heatmap and a normal map to recover the 6-DoF grasping poses represented by our novel representation parameterized with 2D keypoints with corresponding depth, grasping direction, grasping width, and angle. Extensive experiments in real scenes demonstrate that our method can achieve competitive results in grasping common objects and surpass the depth-based competitor by a large margin in grasping photometrically challenging objects. To further stimulate robotic manipulation research, we annotate and open-source a multi-view grasping dataset in the real world containing 44 sequence collections of mixed photometric complexity with nearly 20M accurate grasping labels. Guangyao Zhai, Dianye Huang, Yan Di, Fabian Manhardt, Federico Tombari, Nassir Navab, Benjamin Busam |
ICRA | 1 |
| 2023 | CommonScenes: Generating Commonsense 3D Indoor Scenes with Scene Graphs
Guangyao Zhai, Evin Pinar Örnek, Yan Di, Federico Tombari, Nassir Navab, Benjamin Busam |
NeurIPS | 1 |
| 2023 | DDF-HO: Hand-Held Object Reconstruction via Conditional Directed Distance FieldabstractReconstructing hand-held objects from a single RGB image is an important and challenging problem. Existing works utilizing Signed Distance Fields (SDF) reveal limitations in comprehensively capturing the complex hand-object interactions, since SDF is only reliable within the proximity of the target, and hence, infeasible to simultaneously encode local hand and object cues. To address this issue, we propose DDF-HO, a novel approach leveraging Directed Distance Field (DDF) as the shape representation. Unlike SDF, DDF maps a ray in 3D space, consisting of an origin and a direction, to corresponding DDF values, including a binary visibility signal determining whether the ray intersects the objects and a distance value measuring the distance from origin to target in the given direction. We randomly sample multiple rays and collect local to global geometric features for them by introducing a novel 2D ray-based feature aggregation scheme and a 3D intersection-aware hand pose embedding, combining 2D-3D features to model hand-object interactions. Extensive experiments on synthetic and real-world datasets demonstrate that DDF-HO consistently outperforms all baseline methods by a large margin, especially under Chamfer Distance, with about 80% leap forward. Codes are available at https://github.com/ZhangCYG/DDFHO. Chenyangguang Zhang, Yan Di, Ruida Zhang, Guangyao Zhai, Fabian Manhardt, Federico Tombari, Xiangyang Ji |
NeurIPS | 4 |
| 2020 | Semantic Graph Based Place Recognition for 3D Point CloudsabstractDue to the difficulty in generating the effective descriptors which are robust to occlusion and viewpoint changes, place recognition for 3D point cloud remains an open issue. Unlike most of the existing methods that focus on extracting local, global, and statistical features of raw point clouds, our method aims at the semantic level that can be superior in terms of robustness to environmental changes. Inspired by the perspective of humans, who recognize scenes through identifying semantic objects and capturing their relations, this paper presents a novel semantic graph based approach for place recognition. First, we propose a novel semantic graph representation for the point cloud scenes by reserving the semantic and topological information of the raw point cloud. Thus, place recognition is modeled as a graph matching problem. Then we design a fast and effective graph similarity network to compute the similarity. Exhaustive evaluations on the KITTI dataset show that our approach is robust to the occlusion as well as viewpoint changes and outperforms the state-of-the-art methods with a large margin. Our code is available at: https://github.com/kxhit/SG_PR. Xin Kong, Xuemeng Yang, Guangyao Zhai, Xiangrui Zhao, Xianfang Zeng, Mengmeng Wang 0005, Yong Liu 0007, Wanlong Li |
IROS | 3 |
| 2020 | PoseConvGRU: A Monocular Approach for Visual Ego-motion Estimation by Learning
Guangyao Zhai, Liang Liu 0007, Linjian Zhang, Yong Liu 0007, Yunliang Jiang |
Pattern Recognit. | 1 |
| 2019 | Unsupervised Learning of Scene Flow Estimation Fusing with Local RigidityabstractScene flow estimation in the dynamic scene remains a challenging task. Computing scene flow by a combination of 2D optical flow and depth has shown to be considerably faster with acceptable performance. In this work, we present a unified framework for joint unsupervised learning of stereo depth and optical flow with explicit local rigidity to estimate scene flow. We estimate camera motion directly by a Perspective-n-Point method from the optical flow and depth predictions, with RANSAC outlier rejection scheme. In order to disambiguate the object motion and the camera motion in the scene, we distinguish the rigid region by the re-project error and the photometric similarity. By joint learning with the local rigidity, both depth and optical networks can be refined. This framework boosts all four tasks: depth, optical flow, camera motion estimation, and object motion segmentation. Through the evaluation on the KITTI benchmark, we show that the proposed framework achieves state-of-the-art results amongst unsupervised methods. Our models and code are available at https://github.com/lliuz/unrigidflow. Liang Liu 0007, Guangyao Zhai, Wenlong Ye, Yong Liu 0007 |
IJCAI | 2 |
| 2019 | PASS3D: Precise and Accelerated Semantic Segmentation for 3D Point CloudabstractIn this paper, we propose PASS3D to achieve point-wise semantic segmentation for 3D point cloud. Our framework combines the efficiency of traditional geometric methods with robustness of deep learning methods, consisting of two stages: At stage -1, our accelerated cluster proposal algorithm will generate refined cluster proposals by segmenting point clouds without ground, capable of generating less redundant proposals with higher recall in an extremely short time; stage -2 we will amplify and further process these proposals by a neural network to estimate semantic label for each point and meanwhile propose a novel data augmentation method to enhance the network's recognition capability for all categories especially for non-rigid objects. Evaluated on KITTI raw dataset, PASS3D stands out against the state-of-the-art on some results, making itself competent to 3D perception in autonomous driving system. Our source code will be open-sourced. A video demonstration is available at https://www.youtube.com/watch?v=cukEqDuP_Qw. Xin Kong, Guangyao Zhai, Baoquan Zhong, Yong Liu 0007 |
IROS | 2 |
| 2018 | Millimeter Wave Radar Target Tracking Based on Adaptive Kalman FilterabstractWith the continuous development of the intelligent transportation industry, target tracking has become an important research direction. Under normal circumstances, due to the complex road environment and changing backgrounds, millimeter wave radar has more interference when detecting targets. In addition to the variety of targets in the road and the different scattering intensity of multiple parts, the interference of the flicker noise on the radar must be considered. The combination of these noises can affect the accuracy of radar measurement and even make the radar to lose the target for a short time. The paper constructs a target tracking model based on adaptive Sage-Husa Kalman filter algorithm to track radar signals. The algorithm can not only estimate the real-time state of the system, but also estimate and modify the parameters of the system and the statistical parameters of the noise, so that the system model is closer to the current real state of the system, thus improving the accuracy of the target tracking. Even if radar loses its target in a short time, the target tracking model can estimate the approximate value of the true value of the target. The experimental results show that this method can track the radar target accurately and estimate the position information of the lost target. Guangyao Zhai, Cheng Wu 0001, Yiming Wang 0003 |
Intelligent Vehicles Symposium | 1 |