VLDB 2026 Research / reviewers in the wild / expert
Hu Cheng
dblp:29/3720
· DBLP profile ↗
18ranked-venue papers
7as first author
14since 2021 · last 2026
0000-0002-2090-1362ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 3 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 3 first-author · 5 since 2021Systems, architecture and hardware · 5 · 3 first-author · 2 since 2021Computer networks · 4 · 1 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Vision-based single-stage grasp pose estimator with rotated anchors and automatic label generation
Hu Cheng, Yingying Wang 0003, Max Q.-H. Meng |
Sci. China Inf. Sci. | 1 |
| 2026 | UPDet: Learning-Based Darting-Out Pedestrian Detection by Ultrasonic SensorsabstractUltrasonic sensors are among the critical components in vehicular applications, primarily due to their cost-effectiveness, which is also the main constraint limiting the performance of vehicle-embedded ultrasonic sensors, particularly in detecting dynamic objects. This paper introduces UPDet, a benchmark system for detecting and localizing pedestrians darting out in front of vehicles using only low-cost ultrasonic sensors. The system embeds a novel prototype system for simultaneous signal collection of raw ultrasonic echoes and automatic labeling sensor data, including RGB images and 3D LiDAR point clouds. During the offline signal processing phase, the 3D positions of pedestrians, detected from the fusion of RGB and LiDAR data, are used to annotate the echo data captured by the ultrasonic sensors. Various models that are constructed by 2D and 1D layers with specially designed attention mechanisms are trained and evaluated. More than 13 hours of synchronized sensor data in urban environments are collected. The trained model achieved an 82.28% detection accuracy on the split test set. The high detection performance of the UPDet system, when embedded in the NVIDIA Xavier platform, further demonstrates the substantial potential of ultrasonic-based sensory perception for enhancing automotive safety. Yingying Wang 0003, Chenyu Jin, Hu Cheng |
IEEE Internet Things J. | 3 |
| 2026 | Mixgaze: a dually supervised mixed attention network for gaze estimation
Ziyang Wu, Yin Lin, Hu Cheng, Caihua Kong, Wengang Zhou 0001, Houqiang Li |
Multim. Syst. | 3 |
| 2026 | Automated Action Generation Based on Action Field for Robotic Garment Smoothing and AlignmentabstractGarment manipulation using robotic systems is a challenging task due to the diverse shapes and deformable nature of fabric. In this paper, we propose a novel method for robotic garment smoothing and alignment that significantly improves the accuracy while reducing computational time compared to previous approaches. Our method features an action generator that directly interprets scene images and generates pixel-wise end-effector action vectors using a neural network. The network also predicts a manipulation score map that ranks potential actions, allowing the system to select the most effective action. Extensive simulation experiments demonstrate that our method achieves higher smoothing and alignment performances and faster computation time than previous approaches. Real-world experiments show that the proposed method generalizes well to different garment types and successfully flattens garments. Hu Cheng, Fuyuki Tokuda, Kazuhiro Kosuge |
IEEE Trans Autom. Sci. Eng. | 1 |
| 2025 | Exploratory Analysis of Brainstem fMRI Data During Sustained Phonation
Carey Smith, Hu Cheng, Pertti Palo, Daniel Aalto, Steven M. Lulich |
INTERSPEECH | 2 |
| 2025 | A Learning-Based Sequence-to-Sequence WiFi Fingerprinting Framework for Accurate Pedestrian Indoor Localization Using Unconstrained RSSIabstractIndoor location-based services are essential to daily life but lack a standard like outdoor Global Positioning System (GPS). WiFi received signal strength (RSS) is an optimal option thanks to its ubiquitous deployment, though existing research typically uses only a few constrained devices. We present a learning-based WiFi RSS indicator (RSSI) fingerprinting method designed for general environments. RSSI samples are collected in daily scenarios with hundreds of WiFi access points (APs) without disclosing their locations, and with randomly distributed reference points. The continuously measured RSSI values across multiple timestamps are treated as sequences, serving as fingerprints corresponding to location sequences. Instead of using all detected APs, we discard those sensed only in limited timestamps, i.e., APs with confined coverage and limited identification positions. We then employ various 1D feature extractors to estimate the location sequence from the refined RSSI indices. Our method outperforms state-of-the-art methods on open-access datasets from a small office with densely equipped WiFi APs and a larger university campus with sparse WiFi signals. Real-world experiments on the CUHK campus further demonstrate the statistical consistency of the proposed method. We share the data collection code and self-collected data to facilitate future studies. Yingying Wang 0003, Hu Cheng, Max Q.-H. Meng |
IEEE Internet Things J. | 2 |
| 2025 | Faster-PGYOLO: an efficient framework for floating debris detection in inland waters
Hongru Wang 0009, Hu Cheng |
Vis. Comput. | 2 |
| 2024 | Anchor-Based Multi-Scale Deep Grasp Pose Detector With Encoded Angle RegressionabstractAn intelligent robot grasping system should be able to automatically grasp a variety of objects that have never been seen, which requires accurate and efficient grasp pose detection. To this end, we propose a deep grasp detector designed for the robot equipped with a parallel gripper. The deep model consumes RGB or depth data and extracts features via a feature pyramid network (FPN), followed by multiple grasp prediction units to output grasp parameters in a single stage without refining process. Attaching grasp prediction units to different FPN stages increases the model capability to predict different-size grasps. Furthermore, in each prediction unit, the grasp parameters are regressed with the horizontal anchor as a reference to overcome the challenges posed by the various shapes of the grasp regions. We improve the accuracy and efficiency of grasp rotation estimation by regressing the angle directly and encoding the angle with a continuous Gaussian-like curve during training. This encoded angle regression strategy provides distance information of different angle predictions without introducing additional computational costs. Evaluations on three datasets prove the superior performance of our method than state of the arts. The experiments in real scenarios further validate the effectiveness of our grasping system.Note to Practitioners—This paper proposes a robot system that can automatically grasp novel objects with a parallel gripper and RGB-D camera. We focus on generating accurate grasp configurations for various objects using the captured color or depth image, which is the cornerstone of a successful grasp. To obtain effective and efficient grasp pose detection, we present a deep model that generates robust grasp poses represented by rotated bounding boxes for multiple novel objects. The first step of the grasp detector is to capture the image features through a feature pyramid network (FPN). Then, we attach separate grasp prediction units to each layer of the FPN stage and adopt anchors as references to make the model robust to variable grasp rectangle sizes. In each grasp prediction unit, two separate subnetworks are used to directly output the grasp rectangles and their probabilities, without using an extra second stage to refine the predicted grasp areas. For the prediction of rotation angle, we encode the rectangle angles with a continuous Gaussian-like curve during training to improve the prediction accuracy. Our grasp detector is trained and tested on three datasets and validated on real-scene grasp experiments. Comparisons with state-of-the-art methods show that our model is more accurate while maintaining high efficiency. The proposed grasp detection model can be applied to generate stable grasps for novel objects with different shapes, colors, and materials. Our grasping system is capable of working in multiple scenarios, including homes, factories, and warehouses. Hu Cheng, Yingying Wang 0003, Max Q.-H. Meng |
IEEE Trans Autom. Sci. Eng. | 1 |
| 2024 | Learning to Reorient Objects With Stable Placements Afforded by Extrinsic SupportsabstractReorienting objects by using supports is a practical yet challenging manipulation task. Owing to the intricate geometry of objects and the constrained feasible motions of the robot, multiple manipulation steps are required for object reorientation. In this work, we propose a pipeline for predicting various object placements from point clouds. This pipeline comprises three stages: a pose generation stage, followed by a pose refinement stage, and culminating in a placement classification stage. We also propose an algorithm to construct manipulation graphs based on point clouds. Feasible manipulation sequences are determined for the robot to transfer object placements. Both simulated and real-world experiments demonstrate that our approach is effective. The simulation results underscore our pipeline’s capacity to generalize to novel objects in random start poses. Our predicted placements exhibit a 20% enhancement in accuracy compared to the state-of-the-art baseline. Furthermore, the robot finds feasible sequential steps in the manipulation graphs constructed by our algorithm to accomplish object reorientation manipulation.Note to Practitioners—Object reorientation is a prevalent manipulation task in both domestic and industrial manufacturing scenarios. Extrinsic supporting items are often used to provide diverse object placements that allow for feasible grasp configurations for robotic manipulation. In previous methods, utilizing mesh models of objects was necessary to ascertain stable placements and construct manipulation graphs. In this work, we propose a data-driven approach to predict various object placements conditioned on point clouds. Moreover, we use predicted point cloud placements to construct manipulation graphs, which facilitate collision-free pick-and-place steps to reorient objects. Our approach demonstrates the capacity to generalize to novel objects. In future work, we will enhance the performance of our pipeline by optimizing the distance metric used for measuring pose discrepancies and improving the classifier model. Peng Xu 0006, Hu Cheng, Jiankun Wang 0001, Max Q.-H. Meng |
IEEE Trans Autom. Sci. Eng. | 2 |
| 2023 | Understanding the Gain of Deploying IRSs in Large-Scale Heterogeneous Cellular NetworksabstractAs the superior improvement on wireless network coverage, spectrum efficiency and energy efficiency, Intelligent reflecting surface (IRS) has received more and more attention. In this work, we consider a large-scale IRS-assisted heterogeneous cellular network (HCN) consisting of$K\ (K\geq 2)$tiers of base stations (BSs) and one tier of passive IRSs. With tools from stochastic geometry, we analyze the coverage probability and network spatial throughput of the downlink IRS-assisted$K$-tier HCN. Compared with the conventional HCN, we observe the significant gain achieved by IRSs in coverage probability and network spatial throughput. The proposed analytical framework can be used to understand the limit of gain achieved by IRSs in HCN. Hu Cheng, Linyi Zhang, Jiahui Li 0002, Xijun Wang 0001, Tony Q. S. Quek |
ICC | 1 |
| 2023 | Spatiotemporal Co-Attention Hybrid Neural Network for Pedestrian Localization Based on 6D IMUabstractIn this paper, we propose spatiotemporal co-attention hybrid neural network (SC-HNN), a novel hybrid neural network model with both spatial and temporal attention mechanisms for pose-invariant inertial odometry. The main idea is to extract both local and global features from a window of IMU measurements for velocity prediction. SC-HNN leverages the convolutional neural network (CNN) to capture the sectional features and long short-term memory (LSTM) recurrent neural network (RNN) to extract the long-range dependencies. Attention mechanisms are designed and embedded in both CNN and LSTM modules for better model representation. Specifically, in the CNN attention block, the convolved features are refined along both channel and element dimensions. For the LSTM module, softmax scoring is applied to update the weights of the hidden states along the temporal axis. We evaluate SC-HNN on the benchmark with the largest and most natural IMU data, RoNIN. Extensive ablation experiments demonstrate the effectiveness of our SC-HNN model. Compared with the state of the art, the 50th percentile accuracy of SC-HNN is 18.21% higher and the 90th percentile accuracy is 21.15% higher for all the phone holders not appeared in the training set. The real scenario inertial tracking trials in the CUHK campus further prove the superior generalization ability of the SC-HNN model. Note to Practitioners—This paper aims at improving the localization accuracy of deep inertial odometry. We focus on the problem of indoor localization only from the low-cost IMU embedded in the smartphone without any restriction on the phone’s daily use. IMU is a perfect solution for indoor localization because of its low power consumption, high privacy protection, and external infrastructure free. This paper suggests a novel hybrid convolutional and recurrent neural network with a set of carefully designed attention mechanisms to improve the representation ability of deep inertial odometry model. Specifically, the convolutional layer is applied to extract the local spatial features among the 6D IMU signals, following a cascaded channel attention module and element attention module to boost the representation ability of CNN. The complex long-term dependencies are then identified by the LSTM layers. To adaptively capture the temporal features of the multimodal inertial signals, an attention mechanism is applied to weigh the hidden states for the generation of the final features. The effectiveness of the SC-HNN design is validated by extensive ablation studies. To the best of our knowledge, our model is the first HNN fused attention mechanism for inertial tracking. Extensive experiments show that the proposed method outperforms the state of the art. Yingying Wang 0003, Hu Cheng, Max Q.-H. Meng |
IEEE Trans Autom. Sci. Eng. | 2 |
| 2023 | IRS-Assisted RF-Powered IoT Networks: System Modeling and Performance AnalysisabstractEmerged as a promising solution for future wireless communication systems, intelligent reflecting surface (IRS) is capable of reconfiguring the wireless propagation environment by adjusting the phase-shift of a large number of reflecting elements. To quantify the gain achieved by IRSs in the radio frequency (RF) powered Internet of Things (IoT) networks, in this work, we consider an IRS-assisted cellular-based RF-powered IoT network, where the cellular base stations (BSs) broadcast energy signal to IoT devices for energy harvesting (EH) in the charging stage, which is utilized to support the uplink (UL) transmissions in the subsequent UL stage. With tools from stochastic geometry, we first derive the distributions of the average signal power and interference power which are then used to obtain the energy coverage probability, UL coverage probability, overall coverage probability, spatial throughput and power efficiency, respectively. With the proposed analytical framework, we finally evaluate the effect on network performance of key system parameters, such as IRS density, IRS reflecting element number, charging stage ratio, etc. Compared with the conventional RF-powered IoT network, IRS passive beamforming brings the same level of enhancement in both energy coverage and UL coverage, leading to the unchanged optimal charging stage ratio when maximizing spatial throughput. Zelun Zhao, Hu Cheng, Jiangbin Lyu, Xijun Wang 0001, Yan Zhang 0006, Tony Q. S. Quek |
IEEE Trans. Commun. | 3 |
| 2022 | A2DIO: Attention-Driven Deep Inertial Odometry for Pedestrian Localization based on 6D IMUabstractIn this work, we propose A2DIO, a novel hybrid neural network model with a set of carefully designed attention mechanisms for pose invariant inertial odometry. The key idea is to extract both local and global features from the window of IMU measurements for velocity prediction. A2DIO leverages the convolutional neural network (CNN) to capture the sectional features and long-short term memory (LSTM) recurrent neural network to extract long-range dependencies. In both CNN and LSTM modules, attention mechanisms are designed and embedded for better model representation. Specifically, in the CNN attention block, the convolved features are refined along both channel and spatial dimensions, respectively. For the LSTM module, softmax scoring is applied to update the weights of the hidden states along the temporal axis. We evaluate A2DIO on the benchmark with the largest and most natural IMU data, RoNIN. Extensive ablation experiments demonstrate the effectiveness of our A2DIO model. Compared with the state of the art, the 50th percentile accuracy of A2DIO is 18.21 % higher and the 90th percentile accuracy is 21.15 % higher for all the phone holders not appeared in the training set. Yingying Wang 0003, Hu Cheng, Max Q.-H. Meng |
ICRA | 2 |
| 2021 | Grasp Pose Detection from a Single RGB ImageabstractGrasp pose detection generates the position and orientation of the robot end-effector to grasp objects from the RGB or RGB-D image. In this paper, we propose a novel grasp pose detection network that generates 3-DOF grasp poses using the RGB image. The network follows the anchor-based object detection pipeline and incorporates the angle detection unit. Furthermore, we redesign the grasp angle predictor with a classification unit to increase the accuracy of grasp pose rotation estimation. Our method classifies the prediction angle densely in contrast with the previous regression method or sparse classification method. Moreover, an angle smooth label is designed to avoid the sudden change of the angle regression loss caused by the periodic property of the angle. We validate our algorithm on Cornell Grasp Dataset and obtain a higher detection accuracy than the state-of-the-art method. The real scenario experiment also proves the effectiveness of our method. The robot equipped with the parallel gripper achieves a 96.4% grasp success rate. Hu Cheng, Yingying Wang 0003, Max Q.-H. Meng |
IROS | 1 |
| 2020 | High Accuracy and Efficiency Grasp Pose Detection Scheme with Dense PredictionsabstractLearning-based grasp pose detection algorithms have boosted the performance of robot grasping, but they usually need manually fine-tuning steps to find the balance between detection accuracy and efficient. In this paper, we discard these intermediate procedures, like sampling grasps and generating grasp proposals, and propose an end-to-end grasp pose detection model. Our model uses the RGB image as the input and predicts the single grasp pose in each small grid of the image. Furthermore, the best grasps are found by non-maximum suppression (NMS) strategy. The clustering and ranking procedures are left for NMS while the network only generates dense grasp predictions, which keeps the network simple and efficient. To achieve dense predictions, the predicted grasps of our detection model are represented by the 6 channels images with each pixel location representing a rated grasp. To the best of our knowledge, our model is the first neural network that attaches a grasp pose in pixel level. The model achieves 96.5% accuracy which costs 14ms for prediction of a 480×360 resolution RGB image in Cornell Grasp Dataset, and 90.4% robot grasping success rate for unknown objects with a parallel plate gripper in the real environment. Hu Cheng, Danny Ho, Max Q.-H. Meng |
ICRA | 1 |
| 2020 | Real-Time Robot End-Effector Pose Estimation with Deep NetworkabstractIn this paper, we propose a novel algorithm that estimates the pose of the robot end effector using depth vision. The input to our system is the segmented robot hand point cloud from a depth sensor. Then a neural network takes a point cloud as input and outputs the position and orientation of the robot end effector in the camera frame. The estimated pose can serve as the input of the controller of the robot to reach a specific pose in the camera frame. The training process of the neural network takes the simulated rendered point cloud generated from different poses of the robot hand mesh. At test time, one estimation of a single robot hand pose is reduced to 10ms on gpu and 14ms on cpu, which makes it suitable for close loop robot control system that requires to estimate hand pose in an online fashion. We design a robot hand pose estimation experiment to validate the effectiveness of our algorithm working in the real situation. The platform we used includes a Kinova Jaco 2 robot arm and a Kinect v2 depth sensor. We describe all the processes that use vision to improve the accuracy of pose estimation of the robot end-effector. We demonstrate the possibility of using point cloud to directly estimate the robot's end-effector pose and incorporate the estimated pose into the controller design of the robot arm. Hu Cheng, Yingying Wang 0003, Max Q.-H. Meng |
IROS | 1 |
| 2020 | Pedestrian Motion Tracking by Using Inertial Sensors on the SmartphoneabstractInertial Measurement Unit (IMU) has long been a dream for stable and reliable motion estimation, especially in indoor environments where GPS strength limits. In this paper, we propose a novel method for position and orientation estimation of a moving object only from a sequence of IMU signals collected from the phone. Our main observation is that human motion is monotonous and periodic. We adopt the Extended Kalman Filter and use the learning-based method to dynamically update the measurement noise of the filter. Our pedestrian motion tracking system intends to accurately estimate planar position, velocity, heading direction without restricting the phone's daily use. The method is not only tested on the self-collected signals, but also provides accurate position and velocity estimations on the public RIDI dataset, i.e., the absolute transmit error is 1.28m for a 59-second sequence. Yingying Wang 0003, Hu Cheng, Max Q.-H. Meng |
IROS | 2 |
| 2000 | Design and Implementation of Java Just-in-Time Compiler
Jia Mei, Hu Cheng |
J. Comput. Sci. Technol. | 3 |