EDBT 2026 Demo / reviewers in the wild / expert
Yanzi Miao
dblp:61/550
· DBLP profile ↗
14ranked-venue papers
1as first author
13since 2021 · last 2026
0000-0002-2688-7477ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 7 · 7 since 2021Artificial intelligence and machine learning · 6 · 1 first-author · 5 since 2021Systems, architecture and hardware · 4 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SDAD: Structured Semantic Disentanglement and Attention Diffusion for Open-Vocabulary GraspingabstractOpen-vocabulary grasping is essential for enabling robots to operate robustly in open and dynamic real-world environments. Current methods typically focus on predicting grasp poses from fused cross-modal features. The coupling of visual features and textual semantics in the feature space results in scene-level representations. However, such representations lack object-level discrimination, thereby limiting the model’s ability to recognize and generalize to novel targets. Moreover, existing positional encoding methods employed in cross-modal modeling lack the capacity to effectively capture spatial relationships between regions, which limits the model’s spatial reasoning performance. To tackle these challenges, we propose a method that combines Structured Semantic Disentanglement with Attention Diffusion (SDAD) for open-vocabulary robotic grasping. Specifically, to improve the semantic disentanglement in cross-modal features, a language-guided Disentanglement strategy that regularizes the feature space is proposed to disentangle target semantics features and conditional semantics features. To further improve spatial context modeling, we introduce an attention diffusion mechanism inspired by Fick’s law, which describes natural diffusion driven by concentration gradients, enabling attention to propagate smoothly from condition-feature-anchored regions across the scene. The target features are subsequently used to complete the matching and localization of the corresponding object regions. Extensive quantitative and qualitative experiments demonstrate that our method outperforms existing approaches in open-vocabulary grasping tasks. Moreover, in real-world scenarios, the proposed model achieves a grasp success rate of 84% on base categories and 66% on novel categories, outperforming the state-of-the-art method GLIPv2 by 7% and 5%, respectively. Jin Liu 0018, Yixuan Zhou 0003, Yunfeng Kang, Yanzi Miao, Hesheng Wang 0001 |
IEEE Trans Autom. Sci. Eng. | 5 |
| 2026 | HandMLP: A GAT-Enhanced MLP Network for Robust 3-D Hand Pose Estimationabstract3-D hand pose estimation plays a crucial role in various applications, including virtual interaction, augmented reality, and human–robot collaboration. However, the complex articulation of hand joints and frequent self-occlusions make this task highly challenging. Existing methods based on convolutional or graph architectures are effective for local modeling, but often struggle to jointly capture global context and structured joint dependencies. To address this limitation, we proposeHandMLP, a novel framework that integrates multilayer perceptron (MLP)-based global modeling with graph attention for structured joint reasoning. Specifically, the MLP module enhances global semantic representation and nonlinear fitting capacity, while the graph attention module explicitly models local structural relationships and dynamic dependencies among hand joints. In addition, we introduce a voting-based initialization strategy to generate coarse pose hypotheses, which are progressively refined through collaborative reasoning, improving robustness under sparse or noisy point clouds. Extensive experiments on three benchmark datasets demonstrate thatHandMLPconsistently outperforms state-of-the-art methods, achieving superior accuracy and robustness in 3-D hand pose estimation. Changbo Gao, Yin-Dong Zheng, Mingxi Zhuang, Yanzi Miao, Hesheng Wang 0001 |
IEEE Trans. Ind. Informatics | 5 |
| 2026 | Non-Communicative Decentralized Cooperative Navigation: Reinforcement Learning From Point CloudsabstractLast-mile transport is costly and operationally fragile due to frequent layout changes, mixed traffic with pedestrians and other robots, and unreliable connectivity that makes precise maps hard to maintain. Classical multi-agent planners rely on globally consistent maps or communication. Local BEV pipelines depend on tightly registered perception stacks. End-to-end vision models can be brittle across scene changes. While 3D LiDAR offers geometry that transfers well, its high dimensionality and occlusions complicate real-time, multi-agent control without messaging. We propose an end-to-end mapless, communication-free navigator that consumes a single onboard 3D LiDAR scan and outputs discrete actions for decentralized execution. The approach introduces a dual-channel LiDAR projection that preserves near-field free space while aligning goal guidance. A graph-attention interaction head infers neighbors’ motion tendencies from local observations only, which guides different coordination strategies for pedestrians and vehicles. A structure-aware dense reward stabilizes learning around occlusion-prone layouts (e.g., long walls, U-shaped bays). Our simulations demonstrate strong generalizability and scalability across layouts and sizes. The module trained only in random scenes with 10 agents transfers to crowd and warehouse styles, which to clusters with over$10^{3}$agents. Real-world trials validate sim-to-real transfer without retraining. We deploy the model on transport vehicles and small robots under changing environments. The results indicate robust, scalable, and deployment-friendly navigation for last-mile operations. Zhe Liu 0022, Yanzi Miao, Hesheng Wang 0001 |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2025 | FreeDriveRF: Monocular RGB Dynamic NeRF Without Poses for Autonomous Driving via Point-Level Dynamic-Static DecouplingabstractDynamic scene reconstruction for autonomous driving enables vehicles to perceive and interpret complex scene changes more precisely. Dynamic Neural Radiance Fields (NeRFs) have recently shown promising capability in scene modeling. However, many existing methods rely heavily on accurate poses inputs and multi-sensor data, leading to increased system complexity. To address this, we propose FreeDriveRF, which reconstructs dynamic driving scenes using only sequential RGB images without requiring poses inputs. We innovatively decouple dynamic and static parts at the early sampling level using semantic supervision, mitigating image blurring and artifacts. To overcome the challenges posed by object motion and occlusion in monocular camera, we introduce a warped ray-guided dynamic object rendering consistency loss, utilizing optical flow to better constrain the dynamic modeling process. Additionally, we incorporate estimated dynamic flow to constrain the pose optimization process, improving the stability and accuracy of unbounded scene reconstruction. Extensive experiments conducted on the KITTI and Waymo datasets demonstrate the superior performance of our method in dynamic scene modeling for autonomous driving. Our implementation will be available at https://github.com/IRMVLab/FreeDriveRF. Yue Wen, Siting Zhu 0001, Yanzi Miao, Hesheng Wang 0001 |
ICRA | 5 |
| 2025 | RL-GSBridge: 3D Gaussian Splatting Based Real2Sim2Real Method for Robotic Manipulation LearningabstractSim-to-Real refers to the process of transferring policies learned in simulation to the real world, which is crucial for achieving practical robotics applications. However, recent Sim2real methods either rely on a large amount of augmented data or large learning models, which is inefficient for specific tasks. In recent years, with the emergence of radiance field reconstruction methods, especially 3D Gaussian splatting, it has become possible to construct realistic real-world scenes. To this end, we propose RL-GSBridge, a novel real-to-sim-to-real framework which incorporates 3D Gaussian Splatting into the conventional RL simulation pipeline, enabling zero-shot sim-to-real transfer for vision-based deep reinforcement learning. We introduce a mesh-based 3D GS method with soft binding constraints, enhancing the rendering quality of mesh models. Then utilizing a GS editing approach to synchronize the rendering with the physics simulator, RL-GSBridge could reflect the visual interactions of the physical robot accurately. Through a series of sim-to-real experiments, including grasping and pick-and-place tasks, we demonstrate that RL-GSBridge maintains a satisfactory success rate in real-world task completion during sim-to-real transfer. Furthermore, a series of rendering metrics and visualization results indicate that our proposed mesh-based 3D GS reduces artifacts in unstructured objects, demonstrating more realistic rendering performance. Guangming Wang 0001, Yanzi Miao, Fan Xu 0004, Hesheng Wang 0001 |
ICRA | 5 |
| 2025 | MovSAM: A Single-image Moving Object Segmentation Framework Based on Deep ThinkingabstractMoving object segmentation plays a vital role in understanding dynamic visual environments. While existing methods rely on multi-frame image sequences to identify moving objects, single-image MOS is critical for applications like motion intention prediction and handling camera frame drops. However, segmenting moving objects from a single image remains challenging for existing methods due to the absence of temporal cues. To address this gap, we propose MovSAM, the first framework for single-image moving object segmentation. MovSAM leverages a Multimodal Large Language Model (MLLM) enhanced with Chain-of-Thought (CoT) prompting to search the moving object and generate text prompts based on deep thinking for segmentation. These prompts are cross-fused with visual features from the Segment Anything Model (SAM) and a Vision-Language Model (VLM), enabling logic-driven moving object segmentation. The segmentation results then undergo a deep thinking refinement loop, allowing MovSAM to iteratively improve its understanding of the scene context and inter-object relationships with logical reasoning. This innovative approach enables MovSAM to segment moving objects in single images by considering scene understanding. We implement MovSAM in the real world to validate its practical application and effectiveness for autonomous driving scenarios where the multi-frame methods fail. Furthermore, despite the inherent advantage of multi-frame methods in utilizing temporal information, MovSAM achieves state-of-the-art performance across public MOS benchmarks, reaching 92.5% on J&F. Our implementation will be available at https://github.com/IRMVLab/MovSAM. Chang Nie, Yiqing Xu, Guangming Wang 0001, Zhe Liu 0022, Yanzi Miao |
IROS | 5 |
| 2024 | Gas-Source Efficiently Active Searching in Unfamiliar EnvironmentsabstractSearching Gas Source actively and efficiently in unknown hazard environments is an important but challenging issue. Using mobile robots to autonomously search and navigate to gas source location provides a promising way. Existing methods are mostly based on the modularization framework which investigates the gas-source search and robot navigation tasks independently, leading to a decoupled approach that results in higher collision risks and lower navigation efficiency. Moreover, existing robot navigation techniques grapple with the intricacies of navigating through unknown environments. To tackle these complexities, we introduce an integrated framework that merges gas source localization with robot navigation. This unified structure, underpinned by an end-to-end learning approach, resolves the inherent conflicts between gas exploration and collision avoidance. Our approach aggregates the local observations (raw 3D-LiDAR data) and the expert guidance information (gas distribution), and directly generates navigation actions by implementing the reinforcement learning with a novel reward function based on region dynamic guidances, thus effectively addressing the challenges of active gas source searching in unknown environments. Simulation results underscore the adaptability of our method to diverse unknown environments, along with its superior gas source searching capabilities compared to conventional approaches. Finally, we conduct real-world experiments to demonstrate our feasibility. Yanzi Miao |
ICRA | 2 |
| 2024 | DnFPlane for Efficient and High-Quality 4D Reconstruction of Deformable Tissues
Ran Bu, Chenwei Xu, Jiwei Shan, Hao Li 0133, Guangming Wang 0001, Yanzi Miao, Hesheng Wang 0001 |
MICCAI (6) | 6 |
| 2023 | 3-D Scene Flow Estimation on Pseudo-LiDAR: Bridging the Gap on Estimating Point Motionabstract3-D scene flow characterizes how the points at the current time flow to the next time in the 3-D Euclidean space, which possesses the capacity to infer autonomously the nonrigid motion of all objects in the scene. The previous methods for estimating scene flow from images have limitations, which split the holistic nature of 3-D scene flow by estimating optical flow and disparity separately. Learning 3-D scene flow from point clouds also faces the difficulties of the gap between synthesized and real data and the sparsity of LiDAR point clouds. In this article, the generated dense depth map is utilized to obtain explicit 3-D coordinates, which achieves direct learning of 3-D scene flow from 2-D images. The stability of the predicted scene flow is improved by introducing the dense nature of 2-D pixels into the 3-D space. Outliers in the generated 3-D point cloud are removed by statistical methods to weaken the impact of noisy points on the 3-D scene flow estimation task. Disparity consistency loss is proposed to achieve more effective unsupervised learning of 3-D scene flow. The proposed method of self-supervised learning of 3-D scene flow on real-world images is compared with a variety of methods for learning on the synthesized dataset and learning on LiDAR point clouds. The comparisons of multiple scene flow metrics are shown to demonstrate the effectiveness and superiority of introducing pseudo-LiDAR point cloud to scene flow estimation. Chaokang Jiang, Guangming Wang 0001, Yanzi Miao, Hesheng Wang 0001 |
IEEE Trans. Ind. Informatics | 3 |
| 2023 | Graph Relational Reinforcement Learning for Mobile Robot Navigation in Large-Scale Crowded EnvironmentsabstractMobile robot autonomous navigation in large-scale environments with crowded dynamic objects and static obstacles is still an essential yet challenging task. Recent works have demonstrated the potential of using deep reinforcement learning to enable autonomous navigation in crowds. However, only considering the human-robot interactions results in short-sighted and unsafe behaviors, and they typically use hand-crafted features and assume the global observation range, leading to large performance declines in large-scale crowded environments. Recent advances have shown the power of graph neural networks to learn local interactions among surrounding objects. In this paper, we consider autonomous navigation task in large-scale environments with crowded static and dynamic objects (such as humans). Particularly, local interactions among dynamic objects are learned for better-understanding their moving tendency and relational graph learning is introduced for aggregating both the object-object interactions and object-robot interactions. In addition, local observations are transformed into graphical inputs to achieve the scalability to various number of surrounding dynamic objects and various static obstacle patterns, and the globally guided reinforcement learning strategy is introduced to achieve the fixed-sized learning model even in large-scale complex environments. Simulation results validate our generalizability to various environments and advanced performance compared with existing works in large-scale crowded environments. In particular, our method with only local observations performs better than the benchmarks with global complete observability. Finally, physical robotic experiments demonstrate our effectiveness and practical applicability in real scenarios. Zhe Liu 0022, Jiaming Li 0011, Guangming Wang 0001, Yanzi Miao, Hesheng Wang 0001 |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2022 | Visual Servoing of Flexible-Link Manipulators by Considering Vibration Suppression Without Deformation MeasurementsabstractVisual servoing and vibration suppression of spatial flexible-link manipulators with a fixed camera setup are addressed in this article. The singular perturbation method is adopted to decouple the dynamic equations of the flexible manipulator; hence, two subsystems that represent the rigid robot motion and flexible-link vibration are obtained, respectively. Then, for the slow subsystem related to the rigid motion, an image-based controller is designed to converge the image errors with the consideration of compensating for the errors of approximating the Jacobian matrix. For the fast subsystem corresponding to the elastic vibration, to eliminate the requirements of measuring the vibration states, an observer is designed to estimate the fast states and then a feedback controller of the fast subsystem is presented to suppress the vibration of the flexible manipulator by using the estimation values. The closed-loop stabilities of the slow and fast subsystem are both proved by employing the Lyapunov theory. Numerical simulations demonstrate the effectiveness of the proposed controller, which shows that the image errors approach zero with the vibration of the flexible manipulator damped out simultaneously. Hesheng Wang 0001, Xinwu Liang, Yanzi Miao |
IEEE Trans. Cybern. | 4 |
| 2022 | Uncalibrated Visual Servoing for a Planar Two Link Rigid-Flexible Manipulator Without Joint-Space-Velocity MeasurementabstractIn this article, to solve trajectory tracing problem and vibration suppression for a planar two-link rigid-flexible manipulator subject to joint-velocity measurement noise, a novel uncalibrated visual servoing control is proposed. To begin with, the manipulator’s dynamic model is established by the assumed mode method (AMM). On this basis, based on the singular perturbation theory, two subsystem controllers are designed, one is slow subsystem controller, and the other one is fast subsystem controller. In the slow subsystem, to cope with the complication of the camera calibration, an adaptive algorithm is formulated to evaluate the parameters of a fixed camera online. Aiming to overcome the challenge that exact joint-velocity measurement may be disturbed by external noise, a nonlinear sliding observer is developed to estimate the state of joint velocity accurately. The asymptotic convergence of image tracking error is proved by means of Lyapunov analysis. Additionally, for the purpose of restraining the flexible beam’s elastic vibration, a linear quadratic regulator (LQR) approach is adopted in the fast subsystem control design. The realistic comparing simulation experiments are presented to demonstrate the performance of the proposed controller. Tian Hao, Hesheng Wang 0001, Fan Xu 0004, Jingchuan Wang, Yanzi Miao |
IEEE Trans. Syst. Man Cybern. Syst. | 5 |
| 2021 | S-VIT: Stereo Visual-Inertial Tracking of Lower Limb for Physiotherapy Rehabilitation in Context of Comprehensive Evaluation of SLAM SystemsabstractEstimation of the motion of an agent and its environment concurrently is done by simultaneous localization and mapping (SLAM). In the recent past, SLAM has made rapid and exciting progress and is used in different fields such as unmanned aerial vehicle (UAV), medical surgeries, and endoscopic procedures. The aim of this article is to devise a more accurate physiotherapy exercise monitoring device on the basis of analysis from eight different SLAM algorithms with criteria including power and memory consumption, CPU heat, and CPU utilization. This article provides a comprehensive evaluation on an embedded platform and is first of its kind, especially providing that SLAM systems ego-motion estimation has never been done so explicitly before. Based on the results of the prior analysis, we proposed a stereo visual-inertial tracking (S-VIT) for lower limb tracking in physiotherapy applications. Our proposed algorithm has significantly improved results compared with the state-of-the-art algorithms. Data sets of various physiotherapy rehabilitation exercises for leg are also collected for detailed validations where the ground truth is acquired with a state-of-the-art motion tracking system, Vicon.Note to Practitioners—Accurate real-time tracking in an unknown environment is a challenging task, especially if high accuracy is needed. In physiotherapy, analyzing the daily recorded data (data acquired from a patient’s body movement) will be beneficial in the process of rehabilitation. However, keeping the daily record of patient body motion during exercise is a difficult task in most of the circumstances, which is due to the nonavailability of precise portable devices for accurate motion tracking. In this article, a solution for tracking the patient’s motion (during physiotherapy) using a small hand-held device is presented. For this purpose, simultaneous localization and mapping (SLAM) is used, which, according to the best of our knowledge, is not used in the field of physiotherapy before. In the first part of this article, we present an extensive analysis of a few SLAM algorithms based on power, memory consumption, CPU heat, and CPU utilization. Based on these results, we select the best algorithm EMoVI-SLAM, and then, we extend the work of this article by modifying EMoVI-SLAM. We propose a new SLAM algorithm called stereo visual-inertial tracking (S-VIT). The proposed algorithm is compared with the EMoVI-SLAM on our data set. We collect the data set of various movements of the lower limb. The result shows that S-VIT outperforms EMoVI-SLAM. Azam Rafique Memon, Hesheng Wang 0001, Yanzi Miao, Xiufeng Zhang |
IEEE Trans Autom. Sci. Eng. | 5 |
| 2008 | Coal and Gas Outburst Prediction Combining a Neural Network with the Dempster-Shafter Evidence
Yanzi Miao, Jianwei Zhang 0001, Houxiang Zhang, Zhongxiang Zhao |
ISNN (2) | 1 |