Yingying Wang 0003

dblp:87/6339-3 · DBLP profile ↗
← Back
13ranked-venue papers
5as first author
11since 2021 · last 2026
0000-0003-3293-0790ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 2 first-author · 4 since 2021Systems, architecture and hardware · 5 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 4 since 2021Computer networks · 3 · 2 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Vision-based single-stage grasp pose estimator with rotated anchors and automatic label generation
Hu Cheng, Yingying Wang 0003, Max Q.-H. Meng
Sci. China Inf. Sci.2
2026 UPDet: Learning-Based Darting-Out Pedestrian Detection by Ultrasonic Sensors
abstract
Ultrasonic sensors are among the critical components in vehicular applications, primarily due to their cost-effectiveness, which is also the main constraint limiting the performance of vehicle-embedded ultrasonic sensors, particularly in detecting dynamic objects. This paper introduces UPDet, a benchmark system for detecting and localizing pedestrians darting out in front of vehicles using only low-cost ultrasonic sensors. The system embeds a novel prototype system for simultaneous signal collection of raw ultrasonic echoes and automatic labeling sensor data, including RGB images and 3D LiDAR point clouds. During the offline signal processing phase, the 3D positions of pedestrians, detected from the fusion of RGB and LiDAR data, are used to annotate the echo data captured by the ultrasonic sensors. Various models that are constructed by 2D and 1D layers with specially designed attention mechanisms are trained and evaluated. More than 13 hours of synchronized sensor data in urban environments are collected. The trained model achieved an 82.28% detection accuracy on the split test set. The high detection performance of the UPDet system, when embedded in the NVIDIA Xavier platform, further demonstrates the substantial potential of ultrasonic-based sensory perception for enhancing automotive safety.
Yingying Wang 0003, Chenyu Jin, Hu Cheng
IEEE Internet Things J.1
2026 EIRM-RL: Epistemic Integrity Risk Monitoring Inspired Safe Reinforcement Learning for Trustworthy Autonomous Navigation
abstract
Reinforcement learning (RL) has shown great potential for autonomous navigation within internet of things (IoT) environments, where various and changing uncertainties pose significant challenges for safe, real-world deployment. Existing safe RL methods typically employ heuristic constraints while neglecting the combined impact of multiple uncertainty sources, reducing robustness and interpretability. Drawing on concepts from global navigation satellite system (GNSS) integrity monitoring, this paper proposes an epistemic integrity risk monitoring reinforcement learning (EIRM-RL) framework to enable trustworthy autonomous navigation under uncertainty. EIRM-RL extends the GNSS protection level concept to RL by utilizing an assembled world model that quantifies and incorporates sensor noise, systematic bias, and epistemic uncertainty. Furthermore, the framework continuously monitors a dynamic epistemic risk probability, which is incorporated into policy optimization as an adaptive safety constraint via Lagrangian duality. This method enables the agent to proactively avoid hazards and effectively balance safety and performance, even in highly uncertain environments. Extensive experiments demonstrate that EIRM-RL achieves superior success rates, collision avoidance, and robustness compared to state-of-the-art safe RL methods, while maintaining high efficiency.
Yingying Wang 0003, Weisong Wen
IEEE Internet Things J.2
2025 MInF: Multi-Band Invariant Feature Learning for Efficient Inertial Navigation
abstract
Neural Inertial Navigation (NIN) plays a pivotal role in self-localization, aiming to infer the position of a mobile entity using noisy data from the onboard inertial measurement unit (IMU). Most existing methods rely on convolutional neural networks (CNNs) to capture dependencies among multiple variables, yet the time-frequency and invariant underlying features of IMU measurements remain underexplored. In this paper, we propose MInF, a Multi-Band Invariant Feature Learning for Efficient Inertial Navigation. The MInF advances mainly in two aspects. First, we design a Wavelet-based Multi-Band Mixer (MBMixer) for neural inertial navigation, which leverages the merits of multi-band 1D wavelet decomposition and Multi-Layer Perceptron (MLP)-based mixing to efficiently extract information in both the time and frequency domains in IMU measurements. Second, we introduce a self-supervised learning (SSL) method for learning invariant underlying features from inertial data without the need for any semantic labels. On the one hand, we learn a multi-task MBMixer via jointly classifying different transformations (i.e., pretext tasks) applied to an input signal for extracting invariant underlying features. On the other hand, we use the learned MBMixer in pretext task as the pre-trained model and fine-tune it to regress velocity in the neural inertial navigation (i.e., downstream task). Extensive experiments conducted on two real-world datasets demonstrate that the proposed MInF achieves SOTA results in neural inertial navigation, leading to 15% performance improvement while maintaining a low memory footprint and computational cost.
Yingying Wang 0003
ECAI2
2025 A Learning-Based Sequence-to-Sequence WiFi Fingerprinting Framework for Accurate Pedestrian Indoor Localization Using Unconstrained RSSI
abstract
Indoor location-based services are essential to daily life but lack a standard like outdoor Global Positioning System (GPS). WiFi received signal strength (RSS) is an optimal option thanks to its ubiquitous deployment, though existing research typically uses only a few constrained devices. We present a learning-based WiFi RSS indicator (RSSI) fingerprinting method designed for general environments. RSSI samples are collected in daily scenarios with hundreds of WiFi access points (APs) without disclosing their locations, and with randomly distributed reference points. The continuously measured RSSI values across multiple timestamps are treated as sequences, serving as fingerprints corresponding to location sequences. Instead of using all detected APs, we discard those sensed only in limited timestamps, i.e., APs with confined coverage and limited identification positions. We then employ various 1D feature extractors to estimate the location sequence from the refined RSSI indices. Our method outperforms state-of-the-art methods on open-access datasets from a small office with densely equipped WiFi APs and a larger university campus with sparse WiFi signals. Real-world experiments on the CUHK campus further demonstrate the statistical consistency of the proposed method. We share the data collection code and self-collected data to facilitate future studies.
Yingying Wang 0003, Hu Cheng, Max Q.-H. Meng
IEEE Internet Things J.1
2025 Learning Safe, Optimal, and Real-Time Flight Interaction With Deep Confidence-Enhanced Reachability Guarantee
abstract
In the low-altitude economy, ensuring the safe and agile flight of unmanned aerial vehicles (UAVs) in dynamic obstacle environments is essential for expanding interactive applications like parcel delivery. While deep reinforcement learning (DRL) shows promise for UAV motion planning and control, its trial-and-error exploration often struggles to ensure both agility and safety, especially under uncertain observational noise. Therefore, this paper proposes a deep confidence-enhanced reachability policy optimization (DCRPO) framework. By integrating safe DRL with nonlinear model predictive control (NMPC), DCRPO achieves high-level safety decisions, complex real-time joint planning and control for UAVs. Furthermore, we develop a deep confidence-enhanced reachability guarantee that constructs a set of stochastically forward-reachable planned trajectories under uncertainty, enabling robust safety collision probability certifications. This safe reachability mechanism adaptively selects belief space actions from planned actions to interact with the environment, further enhancing safety and reducing training time. In extensive experiments of UAVs traversing a fast-moving rectangular gate, the proposed method outperforms other state-of-the-art baseline methods under varying environments in terms of operational robustness. Furthermore, the proposed method significantly reduces overall collision violations and training time, greatly improving both training safety and efficiency. The demonstration video (https://youtu.be/7xkp9U7FSJg) and the source code (https://github.com/ZyyFLY/DCRPO) are also provided.
Yingying Wang 0003, Penggao Yan, Weisong Wen
IEEE Trans. Intell. Transp. Syst.2
2024 Anchor-Based Multi-Scale Deep Grasp Pose Detector With Encoded Angle Regression
abstract
An intelligent robot grasping system should be able to automatically grasp a variety of objects that have never been seen, which requires accurate and efficient grasp pose detection. To this end, we propose a deep grasp detector designed for the robot equipped with a parallel gripper. The deep model consumes RGB or depth data and extracts features via a feature pyramid network (FPN), followed by multiple grasp prediction units to output grasp parameters in a single stage without refining process. Attaching grasp prediction units to different FPN stages increases the model capability to predict different-size grasps. Furthermore, in each prediction unit, the grasp parameters are regressed with the horizontal anchor as a reference to overcome the challenges posed by the various shapes of the grasp regions. We improve the accuracy and efficiency of grasp rotation estimation by regressing the angle directly and encoding the angle with a continuous Gaussian-like curve during training. This encoded angle regression strategy provides distance information of different angle predictions without introducing additional computational costs. Evaluations on three datasets prove the superior performance of our method than state of the arts. The experiments in real scenarios further validate the effectiveness of our grasping system.Note to Practitioners—This paper proposes a robot system that can automatically grasp novel objects with a parallel gripper and RGB-D camera. We focus on generating accurate grasp configurations for various objects using the captured color or depth image, which is the cornerstone of a successful grasp. To obtain effective and efficient grasp pose detection, we present a deep model that generates robust grasp poses represented by rotated bounding boxes for multiple novel objects. The first step of the grasp detector is to capture the image features through a feature pyramid network (FPN). Then, we attach separate grasp prediction units to each layer of the FPN stage and adopt anchors as references to make the model robust to variable grasp rectangle sizes. In each grasp prediction unit, two separate subnetworks are used to directly output the grasp rectangles and their probabilities, without using an extra second stage to refine the predicted grasp areas. For the prediction of rotation angle, we encode the rectangle angles with a continuous Gaussian-like curve during training to improve the prediction accuracy. Our grasp detector is trained and tested on three datasets and validated on real-scene grasp experiments. Comparisons with state-of-the-art methods show that our model is more accurate while maintaining high efficiency. The proposed grasp detection model can be applied to generate stable grasps for novel objects with different shapes, colors, and materials. Our grasping system is capable of working in multiple scenarios, including homes, factories, and warehouses.
Hu Cheng, Yingying Wang 0003, Max Q.-H. Meng
IEEE Trans Autom. Sci. Eng.2
2023 Towards an Accurate Augmented-Reality-Assisted Orthopedic Surgical Robotic System Using Bidirectional Generalized Point Set Registration
abstract
This paper presents a novel augmented reality (AR)-assisted orthopedic surgical robotic system based on Head-Mounted Display (HMD) devices. The proposed system can overlay the preoperative plans over the patient's anatomy and provide useful guidance for surgeons during interventions, with integrated calibration and registration components. A novel bi-directional generalised point set registration algorithm that utilises robust features is developed to accurately align the pre-operative CT and intra-operative patient spaces, which has been demonstrated to outperform existing registration methods. The efficacy of the system is both qualitatively and quantitatively assessed with an in vitro study simulating a total knee arthroplasty (TKA) procedure. The experimental results showed that 1) the system can successfully align the preoperative and intraoperative spaces, with the mean target registration error (TRE) being 2.7771 mm; 2) the models can be properly overlaid to the physical scenarios with the mean AR visualization accuracy being 6.9726 mm.
Zhe Min, Yingying Wang 0003, Max Q.-H. Meng
IROS3
2023 Spatiotemporal Co-Attention Hybrid Neural Network for Pedestrian Localization Based on 6D IMU
abstract
In this paper, we propose spatiotemporal co-attention hybrid neural network (SC-HNN), a novel hybrid neural network model with both spatial and temporal attention mechanisms for pose-invariant inertial odometry. The main idea is to extract both local and global features from a window of IMU measurements for velocity prediction. SC-HNN leverages the convolutional neural network (CNN) to capture the sectional features and long short-term memory (LSTM) recurrent neural network (RNN) to extract the long-range dependencies. Attention mechanisms are designed and embedded in both CNN and LSTM modules for better model representation. Specifically, in the CNN attention block, the convolved features are refined along both channel and element dimensions. For the LSTM module, softmax scoring is applied to update the weights of the hidden states along the temporal axis. We evaluate SC-HNN on the benchmark with the largest and most natural IMU data, RoNIN. Extensive ablation experiments demonstrate the effectiveness of our SC-HNN model. Compared with the state of the art, the 50th percentile accuracy of SC-HNN is 18.21% higher and the 90th percentile accuracy is 21.15% higher for all the phone holders not appeared in the training set. The real scenario inertial tracking trials in the CUHK campus further prove the superior generalization ability of the SC-HNN model. Note to Practitioners—This paper aims at improving the localization accuracy of deep inertial odometry. We focus on the problem of indoor localization only from the low-cost IMU embedded in the smartphone without any restriction on the phone’s daily use. IMU is a perfect solution for indoor localization because of its low power consumption, high privacy protection, and external infrastructure free. This paper suggests a novel hybrid convolutional and recurrent neural network with a set of carefully designed attention mechanisms to improve the representation ability of deep inertial odometry model. Specifically, the convolutional layer is applied to extract the local spatial features among the 6D IMU signals, following a cascaded channel attention module and element attention module to boost the representation ability of CNN. The complex long-term dependencies are then identified by the LSTM layers. To adaptively capture the temporal features of the multimodal inertial signals, an attention mechanism is applied to weigh the hidden states for the generation of the final features. The effectiveness of the SC-HNN design is validated by extensive ablation studies. To the best of our knowledge, our model is the first HNN fused attention mechanism for inertial tracking. Extensive experiments show that the proposed method outperforms the state of the art.
Yingying Wang 0003, Hu Cheng, Max Q.-H. Meng
IEEE Trans Autom. Sci. Eng.1
2022 A2DIO: Attention-Driven Deep Inertial Odometry for Pedestrian Localization based on 6D IMU
abstract
In this work, we propose A2DIO, a novel hybrid neural network model with a set of carefully designed attention mechanisms for pose invariant inertial odometry. The key idea is to extract both local and global features from the window of IMU measurements for velocity prediction. A2DIO leverages the convolutional neural network (CNN) to capture the sectional features and long-short term memory (LSTM) recurrent neural network to extract long-range dependencies. In both CNN and LSTM modules, attention mechanisms are designed and embedded for better model representation. Specifically, in the CNN attention block, the convolved features are refined along both channel and spatial dimensions, respectively. For the LSTM module, softmax scoring is applied to update the weights of the hidden states along the temporal axis. We evaluate A2DIO on the benchmark with the largest and most natural IMU data, RoNIN. Extensive ablation experiments demonstrate the effectiveness of our A2DIO model. Compared with the state of the art, the 50th percentile accuracy of A2DIO is 18.21 % higher and the 90th percentile accuracy is 21.15 % higher for all the phone holders not appeared in the training set.
Yingying Wang 0003, Hu Cheng, Max Q.-H. Meng
ICRA1
2021 Grasp Pose Detection from a Single RGB Image
abstract
Grasp pose detection generates the position and orientation of the robot end-effector to grasp objects from the RGB or RGB-D image. In this paper, we propose a novel grasp pose detection network that generates 3-DOF grasp poses using the RGB image. The network follows the anchor-based object detection pipeline and incorporates the angle detection unit. Furthermore, we redesign the grasp angle predictor with a classification unit to increase the accuracy of grasp pose rotation estimation. Our method classifies the prediction angle densely in contrast with the previous regression method or sparse classification method. Moreover, an angle smooth label is designed to avoid the sudden change of the angle regression loss caused by the periodic property of the angle. We validate our algorithm on Cornell Grasp Dataset and obtain a higher detection accuracy than the state-of-the-art method. The real scenario experiment also proves the effectiveness of our method. The robot equipped with the parallel gripper achieves a 96.4% grasp success rate.
Hu Cheng, Yingying Wang 0003, Max Q.-H. Meng
IROS2
2020 Real-Time Robot End-Effector Pose Estimation with Deep Network
abstract
In this paper, we propose a novel algorithm that estimates the pose of the robot end effector using depth vision. The input to our system is the segmented robot hand point cloud from a depth sensor. Then a neural network takes a point cloud as input and outputs the position and orientation of the robot end effector in the camera frame. The estimated pose can serve as the input of the controller of the robot to reach a specific pose in the camera frame. The training process of the neural network takes the simulated rendered point cloud generated from different poses of the robot hand mesh. At test time, one estimation of a single robot hand pose is reduced to 10ms on gpu and 14ms on cpu, which makes it suitable for close loop robot control system that requires to estimate hand pose in an online fashion. We design a robot hand pose estimation experiment to validate the effectiveness of our algorithm working in the real situation. The platform we used includes a Kinova Jaco 2 robot arm and a Kinect v2 depth sensor. We describe all the processes that use vision to improve the accuracy of pose estimation of the robot end-effector. We demonstrate the possibility of using point cloud to directly estimate the robot's end-effector pose and incorporate the estimated pose into the controller design of the robot arm.
Hu Cheng, Yingying Wang 0003, Max Q.-H. Meng
IROS2
2020 Pedestrian Motion Tracking by Using Inertial Sensors on the Smartphone
abstract
Inertial Measurement Unit (IMU) has long been a dream for stable and reliable motion estimation, especially in indoor environments where GPS strength limits. In this paper, we propose a novel method for position and orientation estimation of a moving object only from a sequence of IMU signals collected from the phone. Our main observation is that human motion is monotonous and periodic. We adopt the Extended Kalman Filter and use the learning-based method to dynamically update the measurement noise of the filter. Our pedestrian motion tracking system intends to accurately estimate planar position, velocity, heading direction without restricting the phone's daily use. The method is not only tested on the self-collected signals, but also provides accurate position and velocity estimations on the public RIDI dataset, i.e., the absolute transmit error is 1.28m for a 59-second sequence.
Yingying Wang 0003, Hu Cheng, Max Q.-H. Meng
IROS1