VLDB 2026 Research / reviewers in the wild / expert
Hao Yu 0010
dblp:64/4832-10
· DBLP profile ↗
13ranked-venue papers
5as first author
9since 2021 · last 2024
0000-0002-2111-6905ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 5 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | MatchU: Matching Unseen Objects for 6D Pose Estimation from RGB-D ImagesabstractRecent learning methods for object pose estimation require resource-intensive training for each individual object instance or category, hampering their scalability in real applications when confronted with previously unseen objects. In this paper, we propose MatchU, a Fuse-Describe-Match strategy for 6D pose estimation from RGB-D images. MatchU is a generic approach that fuses 2D texture and 3D geometric cues for 6D pose prediction of unseen objects. We rely on learning geometric 3D descriptors that are rotation-invariant by design. By encoding pose-agnostic geometry, the learned descriptors naturally generalize to unseen objects and capture symmetries. To tackle ambiguous associations using 3D geometry only, we fuse additional RGB information into our descriptor. This is achieved through a novel attention-based mechanism that fuses cross-modal information, together with a matching loss that leverages the latent space learned from RGB data to guide the descriptor learning process. Extensive experiments reveal the generalizability of both the RGB-D fusion strategy as well as the descriptor efficacy. Benefiting from the novel designs, MatchU surpasses all existing methods by a significant margin in terms of both accuracy and speed, even without the requirement of expensive re-training or rendering. Junwen Huang 0001, Hao Yu 0010, Kuan-Ting Yu, Nassir Navab, Slobodan Ilic, Benjamin Busam |
CVPR | 2 |
| 2024 | RIGA: Rotation-Invariant and Globally-Aware Descriptors for Point Cloud RegistrationabstractSuccessful point cloud registration relies on accurate correspondences established upon powerful descriptors. However, existing neural descriptors either leverage a rotation-variant backbone whose performance declines under large rotations, or encode local geometry that is less distinctive. To address this issue, we introduce RIGA to learn descriptors that are Rotation-Invariant by design and Globally-Aware. From the Point Pair Features (PPFs) of sparse local regions, rotation-invariant local geometry is encoded into geometric descriptors. Global awareness of 3D structures and geometric context is subsequently incorporated, both in a rotation-invariant fashion. More specifically, 3D structures of the whole frame are first represented by our global PPF signatures, from which structural descriptors are learned to help geometric descriptors sense the 3D world beyond local regions. Geometric context from the whole scene is then globally aggregated into descriptors. Finally, the description of sparse regions is interpolated to dense point descriptors, from which correspondences are extracted for registration. To validate our approach, we conduct extensive experiments on both object- and scene-level data. With large rotations, RIGA surpasses the state-of-the-art methods by a margin of 8${}^\circ$in terms of the Relative Rotation Error on ModelNet40 and improves the Feature Matching Recall by at least 5 percentage points on 3DLoMatch. Hao Yu 0010, Ji Hou, Zheng Qin 0002, Mahdi Saleh, Ivan Shugurov, Kai Wang 0037, Benjamin Busam, Slobodan Ilic |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2023 | Deep Graph-based Spatial Consistency for Robust Non-rigid Point Cloud RegistrationabstractWe study the problem of outlier correspondence pruning for non-rigid point cloud registration. In rigid registration, spatial consistency has been a commonly used criterion to discriminate outliers from inliers. It measures the compatibility of two correspondences by the discrepancy between the respective distances in two point clouds. However, spatial consistency no longer holds in non-rigid cases and outlier rejection for non-rigid registration has not been well studied. In this work, we propose Graph-based Spatial Consistency Network (GraphSCNet) to filter outliers for non-rigid registration. Our method is based on the fact that non-rigid deformations are usually locally rigid, or local shape preserving. We first design a local spatial consistency measure over the deformation graph of the point cloud, which evaluates the spatial compatibility only between the correspondences in the vicinity of a graph node. An attention-based non-rigid correspondence embedding module is then devised to learn a robust representation of non-rigid correspondences from local spatial consistency. Despite its simplicity, GraphSCNet effectively improves the quality of the putative correspondences and attains state-of-the-art performance on three challenging benchmarks. Our code and models are available at https://github.com/qinzheng93/GraphSCNet. Zheng Qin 0002, Hao Yu 0010, Yuxing Peng 0001, Kai Xu 0004 |
CVPR | 2 |
| 2023 | Rotation-Invariant Transformer for Point Cloud MatchingabstractThe intrinsic rotation invariance lies at the core of matching point clouds with handcrafted descriptors. However, it is widely despised by recent deep matchers that obtain the rotation invariance extrinsically via data augmentation. As the finite number of augmented rotations can never span the continuous$SO(3)$space, these methods usually show instability when facing rotations that are rarely seen. To this end, we introduce RoITr, a Rotation-Invariant Transformer to cope with the pose variations in the point cloud matching task. We contribute both on the local and global levels. Starting from the local level, we introduce an attention mechanism embedded with Point Pair Feature (PPF)-based coordinates to describe the pose-invariant geometry, upon which a novel attention-based encoder-decoder architecture is constructed. We further propose a global transformer with rotation-invariant cross-frame spatial awareness learned by the self-attention mechanism, which significantly improves the feature distinctiveness and makes the model robust with respect to the low overlap. Experiments are conducted on both the rigid and non-rigid public benchmarks, where RoITr outperforms all the state-of-the-art models by a considerable margin in the low-overlapping scenarios. Especially when the rotations are enlarged on the challenging 3DLoMatch benchmark, RoITr surpasses the existing methods by at least 13 and 5 percentage points in terms of Inlier Ratio and Registration Recall, respectively. Code is publicly available11https://github.com/haoyu94/RoITr. Hao Yu 0010, Zheng Qin 0002, Ji Hou, Mahdi Saleh, Dongsheng Li 0001, Benjamin Busam, Slobodan Ilic |
CVPR | 1 |
| 2023 | Dynamic Hyperbolic Attention Network for Fine Hand-object ReconstructionabstractReconstructing both objects and hands in 3D from a single RGB image is complex. Existing methods rely on manually defined hand-object constraints in Euclidean space, leading to suboptimal feature learning. Compared with Euclidean space, hyperbolic space better preserves the geometric properties of meshes thanks to its exponentially-growing space distance, which amplifies the differences between the features based on similarity. In this work, we propose the first precise hand-object reconstruction method in hyperbolic space, namely Dynamic Hyperbolic Attention Network (DHANet), which leverages intrinsic properties of hyperbolic space to learn representative features. Our method that projects mesh and image features into a unified hyperbolic space includes two modules, i.e. dynamic hyperbolic graph convolution and image-attention hyperbolic graph convolution. With these two modules, our method learns mesh features with rich geometry-image multi-modal information and models better hand-object interaction. Our method provides a promising alternative for fine hand-object reconstruction in hyperbolic space. Extensive experiments on three public datasets demonstrate that our method outperforms most state-of-the-art methods. Zhiying Leng, Mahdi Saleh, Antonio Montanaro, Hao Yu 0010, Yin Wang 0005, Nassir Navab, Xiaohui Liang 0001, Federico Tombari |
ICCV | 5 |
| 2023 | Reduced-Resolution Head for Object Detection
Jiayuan Zhuang, Hao Yu 0010, Xucan Chen |
PRICAI (3) | 3 |
| 2023 | GeoTransformer: Fast and Robust Point Cloud Registration With Geometric TransformerabstractWe study the problem of extracting accurate correspondences for point cloud registration. Recent keypoint-free methods have shown great potential through bypassing the detection of repeatable keypoints which is difficult to do especially in low-overlap scenarios. They seek correspondences over downsampled superpoints, which are then propagated to dense points. Superpoints are matched based on whether their neighboring patches overlap. Such sparse and loose matching requires contextual features capturing the geometric structure of the point clouds. We propose Geometric Transformer, or GeoTransformer for short, to learn geometric feature for robust superpoint matching. It encodes pair-wise distances and triplet-wise angles, making it invariant to rigid transformation and robust in low-overlap cases. The simplistic design attains surprisingly high matching accuracy such that no RANSAC is required in the estimation of alignment transformation, leading to 100 times acceleration. Extensive experiments on rich benchmarks encompassing indoor, outdoor, synthetic, multiway and non-rigid demonstrate the efficacy of GeoTransformer. Notably, our method improves the inlier ratio by 18 ∼ 31 percentage points and the registration recall by over 7 points on the challenging 3DLoMatch benchmark. Zheng Qin 0002, Hao Yu 0010, Yulan Guo, Yuxing Peng 0001, Slobodan Ilic, Dewen Hu, Kai Xu 0004 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2022 | Geometric Transformer for Fast and Robust Point Cloud RegistrationabstractWe study the problem of extracting accurate correspondences for point cloud registration. Recent keypoint-free methods bypass the detection of repeatable keypoints which is difficult in low-overlap scenarios, showing great potential in registration. They seek correspondences over down-sampled superpoints, which are then propagated to dense points. Superpoints are matched based on whether their neighboring patches overlap. Such sparse and loose matching requires contextual features capturing the geometric structure of the point clouds. We propose Geometric Transformer to learn geometric feature for robust superpoint matching. It encodes pair-wise distances and triplet-wise angles, making it robust in low-overlap cases and invariant to rigid transformation. The simplistic design attains surprisingly high matching accuracy such that no RANSAC is required in the estimation of alignment transformation, leading to 100 times acceleration. Our method improves the inlier ratio by 17∼30 percentage points and the registration recall by over 7 points on the challenging 3DLoMatch benchmark. Our code and models are available at https://github.com/qinzheng93/GeoTransformer. Zheng Qin 0002, Hao Yu 0010, Yulan Guo, Yuxing Peng 0001, Kai Xu 0004 |
CVPR | 2 |
| 2021 | CoFiNet: Reliable Coarse-to-fine Correspondences for Robust PointCloud RegistrationabstractWe study the problem of extracting correspondences between a pair of point clouds for registration. For correspondence retrieval, existing works benefit from matching sparse keypoints detected from dense points but usually struggle to guarantee their repeatability. To address this issue, we present CoFiNet - Coarse-to-Fine Network which extracts hierarchical correspondences from coarse to fine without keypoint detection. On a coarse scale and guided by a weighting scheme, our model firstly learns to match down-sampled nodes whose vicinity points share more overlap, which significantly shrinks the search space of a consecutive stage. On a finer scale, node proposals are consecutively expanded to patches that consist of groups of points together with associated descriptors. Point correspondences are then refined from the overlap areas of corresponding patches, by a density-adaptive matching module capable to deal with varying point density. Extensive evaluation of CoFiNet on both indoor and outdoor standard benchmarks shows our superiority over existing methods. Especially on 3DLoMatch where point clouds share less overlap, CoFiNet significantly outperforms state-of-the-art approaches by at least 5% on Registration Recall, with at most two-third of their parameters. Hao Yu 0010, Mahdi Saleh, Benjamin Busam, Slobodan Ilic |
NeurIPS | 1 |
| 2018 | Nominal Data Similarity: A Hierarchical MeasureabstractSimilarity of nominal data plays fundamental roles in numerous fields of both machine learning and data mining. Unlike the similarity of numerical data, that of nominal data is much more difficult to describe, and few efforts have been done for it. Although existing nominal similarity measures can reveal a part of data properties, they suffer from low accuracy due to ignoring value relationships or integrating multi-view relationships inappropriately. In this paper, we propose a novel hierarchical measure for nominal data similarity (HNS). The HNS leverages the intrinsic data characteristics by considering low-level information both within and between attributes, and hierarchically seizes the value distributions, attribute interactions and attribute to object contributions. Meanwhile, it aggregates multi-view relationships trough a bottom to top framework, remaining consistency as well as complementary details. We theoretically analyzed this measure, and experiments on six UCI data sets demonstrate that the HNS outperforms the state-of-the-art nominal similarity measures in term of target alignment and clustering accuracy. Hao Yu 0010, Zhaoning Zhang 0001, Gen Zhang |
IJCNN | 1 |
| 2018 | Merging and Evolution: Improving Convolutional Neural Networks for Mobile ApplicationsabstractCompact neural networks are inclined to exploit “sparsely-connected” convolutions such as depthwise convolution and group convolution for employment in mobile applications. Compared with standard “fully-connected” convolutions, these convolutions are more computationally economical. However, “sparsely-connected” convolutions block the inter-group informa-tion exchange, which induces severe performance degradation. To address this issue, we present two novel operations named merging and evolution to leverage the inter-group information. Our key idea is encoding the inter-group information with a narrow feature map, then combining the generated features with the original network for better representation. Taking advantage of the proposed operations, we then introduce the Merging-and- Evolution (ME) module, an architectural unit specifically designed for compact networks. Finally, we propose a family of compact neural networks called MENet based on ME modules. Extensive experiments on ILSVRC 2012 dataset and PASCAL VOC 2007 dataset demonstrate that MENet consistently outperforms other state-of -the-art compact networks under different computational budgets. For instance, under the computational budget of 140 MFLOPs, MENet surpasses ShuffleNet by 1% and MobileNet by 1.95% on ILSVRC 2012 top-l accuracy, while by 2.3% and 4.1% on PASCAL VOC 2007 mAP, respectively. Zheng Qin 0002, Zhaoning Zhang 0001, Shiqing Zhang, Hao Yu 0010, Jincai Li, Yuxing Peng 0001 |
IJCNN | 4 |
| 2018 | Loss Rank Mining: A General Hard Example Mining Method for Real-time DetectorsabstractModern object detectors usually suffer from low accuracy issues, as foregrounds always drown in tons of back-grounds and become hard examples during training. Compared with those proposal-based ones, real-time detectors are in far more serious trouble since they renounce the use of region-proposing stage which is used to filter a majority of back-grounds for achieving real-time rates. Though foregrounds as hard examples are in urgent need of being mined from tons of backgrounds, a considerable number of state-of-the-art real-time detectors, like YOLO series, have yet to profit from existing hard example mining methods, as using these methods need detectors fit series of prerequisites. In this paper, we propose a general hard example mining method named Loss Rank Mining (LRM) to fill the gap. LRM is a general method for real-time detectors, as it utilizes the final feature map which exists in all real-time detectors to mine hard examples. By using LRM, some elements representing easy examples in final feature map are filtered and detectors are forced to concentrate on hard examples during training. Extensive experiments validate the effectiveness of our method. With our method, the improvements of YOLOv2 detector on auto-driving related dataset KITTI and more general dataset PASCAL VOC are over 5% and 2% mAP, respectively. In addition, LRM is the first hard example mining strategy which could fit YOLOv2 perfectly and make it better applied in series of real scenarios where both real-time rates and accurate detection are strongly demanded. Hao Yu 0010, Zhaoning Zhang 0001, Zheng Qin 0002, Hao Wu 0031, Dongsheng Li 0001, Xicheng Lu |
IJCNN | 1 |
| 2017 | Distribution majorization of corner points by reinforcement learning for moving object detectionabstractCorner points play an important role in moving object detection, especially in the case of free-moving camera. Corner points provide more accurate information than other pixels and reduce the computation which is unnecessary. Previous works only use intensity information to locate the corner points, however, the information that former and the last frames provided also can be used. We utilize the information to focus on more valuable area and ignore the invaluable area. The proposed algorithm is based on reinforcement learning, which regards the detection of corner points as a Markov process. In the Markov model, the video to be detected is regarded as environment, the selections of blocks for one corner point are regarded as actions and the performance of detection is regarded as state. Corner points are assigned to be the blocks which are seperated from original whole image. Experimentally, we select a conventional method which uses marching and Random Sample Consensus algorithm to obtain objects as the main framework and utilize our algorithm to improve the result. The comparison between the conventional method and the same one with our algorithm show that our algorithm reduce 70% of the false detection. Hao Wu 0031, Hao Yu 0010, Dongxiang Zhou, Yongqiang Cheng 0002 |
ICMV | 2 |