Haiyan Wang 0019

dblp:27/59-19 · DBLP profile ↗
← Back
11ranked-venue papers
7as first author
7since 2021 · last 2024
0000-0002-1261-5948ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 9 · 6 first-author · 6 since 2021Artificial intelligence and machine learning · 8 · 5 first-author · 5 since 2021Systems, architecture and hardware · 1
YearPublicationVenuePosition
2024 Sequential Point Clouds: A Survey
abstract
Point clouds have garnered increasing research attention and found numerous practical applications. However, many of these applications, such as autonomous driving and robotic manipulation, rely on sequential point clouds, essentially adding a temporal dimension to the data (i.e., four dimensions) because the information of the static point cloud data could provide is still limited. Recent research efforts have been directed towards enhancing the understanding and utilization of sequential point clouds. This paper offers a comprehensive review of deep learning methods applied to sequential point cloud research, encompassing dynamic flow estimation, object detection & tracking, point cloud segmentation, and point cloud forecasting. This paper further summarizes and compares the quantitative results of the reviewed methods over the public benchmark datasets. Ultimately, the paper concludes by addressing the challenges in current sequential point cloud research and pointing towards promising avenues for future research.
Haiyan Wang 0019, Yingli Tian
IEEE Trans. Pattern Anal. Mach. Intell.1
2022 PSMNet: Position-aware Stereo Merging Network for Room Layout Estimation
abstract
In this paper, we propose a new deep learning-based method for estimating room layout given a pair of 360° panoramas. Our system, called Position-aware Stereo Merging Network or PSMNet, is an end-to-end joint layout-pose estimator. PSMNet consists of a Stereo Pano Pose (SP2) transformer and a novel Cross-Perspective Projection (CP2) layer. The stereo-view SP2 transformer is used to implicitly infer correspondences between views, and can handle noisy poses. The pose-aware CP2layer is designed to render features from the adjacent view to the anchor (reference) view, in order to perform view fusion and estimate the visible layout. Our experiments and analysis validate our method, which significantly outperforms the state-of-the-art layout estimators, especially for large and complex room spaces.
Haiyan Wang 0019, Will Hutchcroft, Yuguang Li, Zhiqiang Wan, Ivaylo Boyadzhiev, Yingli Tian, Sing Bing Kang
CVPR1
2022 Disentangling Object Motion and Occlusion for Unsupervised Multi-frame Monocular Depth
Ziyue Feng, Longlong Jing, Haiyan Wang 0019, Yingli Tian, Bing Li 0008
ECCV (32)4
2022 CoVisPose: Co-visibility Pose Transformer for Wide-Baseline Relative Pose Estimation in 360$^\circ $ Indoor Panoramas
Will Hutchcroft, Yuguang Li, Ivaylo Boyadzhiev, Zhiqiang Wan, Haiyan Wang 0019, Sing Bing Kang
ECCV (32)5
2021 FESTA: Flow Estimation via Spatial-Temporal Attention for Scene Point Clouds
abstract
Scene flow depicts the dynamics of a 3D scene, which is critical for various applications such as autonomous driving, robot navigation, AR/VR, etc. Conventionally, scene flow is estimated from dense/regular RGB video frames. With the development of depth-sensing technologies, precise 3D measurements are available via point clouds which have sparked new research in 3D scene flow. Nevertheless, it remains challenging to extract scene flow from point clouds due to the sparsity and irregularity in typical point cloud sampling patterns. One major issue related to irregular sampling is identified as the randomness during point set abstraction/feature extraction—an elementary process in many flow estimation scenarios. A novel Spatial Abstraction with Attention (SA2) layer is accordingly proposed to alleviate the unstable abstraction problem. Moreover, a Temporal Abstraction with Attention (TA2) layer is proposed to rectify attention in temporal domain, leading to benefits with motions scaled in a larger range. Extensive analysis and experiments verified the motivation and significant performance gains of our method, dubbed as Flow Estimation via Spatial-Temporal Attention (FESTA), when compared to several state-of-the-art benchmarks of scene flow estimation.
Haiyan Wang 0019, Jiahao Pang, Muhammad Asad Lodhi, Yingli Tian, Dong Tian
CVPR1
2021 Subsurface Pipes Detection Using DNN-based Back Projection on GPR Data
abstract
Localization and reconstruction of underground targets, the problem of estimating the position and geometry of the objects from Ground Penetration Radar (GPR), still lies at the core of non-destructive testing (NDT). In this paper, we present MigrationNet, a learning-based approach to detect and visualize subsurface objects. Compared with the existing learning-based method of GPR, our proposed approach could not only detect the hyperbola feature in the raw B-scan image but also interpret hyperbola features into the cross-section image of subsurface pipes. Furthermore, to compare the proposed method with the conventional back-projection methods for GPR data interpretation, a synthetic GPR dataset that mimics the real NDT environment is also introduced in this work. The study indicates the effectiveness of our method, it uses less GPR data for underground pipes reconstruction, produces better GPR imaging results with less computation, and shows the robustness to noise.
Jinglun Feng, Haiyan Wang 0019, Yingli Tian, Jizhong Xiao
WACV3
2021 Self-supervised 4D Spatio-temporal Feature Learning via Order Prediction of Sequential Point Cloud Clips
abstract
Recently 3D scene understanding attracts attention for many applications, however, annotating a vast amount of 3D data for training is usually expensive and time consuming. To alleviate the needs of ground truth, we propose a self-supervised schema to learn 4D spatio-temporal features (i.e. 3 spatial dimensions plus 1 temporal dimension) from dynamic point cloud data by predicting the temporal order of sampled and shuffled point cloud clips. 3D sequential point cloud contains precious geometric and depth information to better recognize activities in 3D space compared to videos. To learn the 4D spatio-temporal features, we introduce 4D convolution neural networks to predict the temporal order on a self-created large scale dataset, NTU-PCLs, derived from the NTU-RGB+D dataset. The efficacy of the learned 4D spatio-temporal features is verified on two tasks: 1) Self-supervised 3D nearest neighbor retrieval; and 2) Self-supervised representation learning transferred for action recognition on a smaller 3D dataset. Our extensive experiments prove the effectiveness of the proposed self-supervised learning method which achieves comparable results w.r.t. the fully-supervised methods on action recognition on MSRAction3D dataset.
Haiyan Wang 0019, Xuejian Rong, Jinglun Feng, Yingli Tian
WACV1
2020 Towards Efficient 3D Point Cloud Scene Completion via Novel Depth View Synthesis
abstract
3D point cloud completion has been a long-standing challenge at scale, and corresponding per-point supervised training strategies suffered from cumbersome annotations. 2D supervision has recently emerged as a promising alternative for 3D tasks, but specific approaches for 3D point cloud completion still remain to be explored. To overcome these limitations, we propose an end-to-end method that directly lifts a single depth map to a completed point cloud. With one depth map as input, a multiway novel depth view synthesis network (NDVNet) is designed to infer coarsely completed depth maps under various viewpoints. Meanwhile, a geometric depth perspective rendering module is introduced to utilize the raw input depth map to generate a re-projected depth map for each view. Therefore, the two parallelly generated depth maps for each view are further concatenated and refined by a depth completion network (DCNet). The final completed point cloud is fused from all refined depth views. Experimental results demonstrate the effectiveness of our proposed approach composed of aforementioned components, to produce high-quality, state-of-the-art results on the popular SUNCG benchmark.
Haiyan Wang 0019, Xuejian Rong, Yingli Tian
ICPR1
2020 GPR-based Subsurface Object Detection and Reconstruction Using Random Motion and DepthNet
abstract
Ground Penetrating Radar (GPR) is one of the most important non-destructive evaluation (NDE) devices to detect the subsurface objects (i.e. rebars, utility pipes) and reveal the underground scene. One of the biggest challenges in GPR based inspection is the subsurface targets reconstruction. In order to address this issue, this paper presents a 3D GPR migration and dielectric prediction system to detect and reconstruct underground targets. This system is composed of three modules: 1) visual inertial fusion (VIF) module to generate the pose information of GPR device, 2) deep neural network module (i.e., DepthNet) which detects B-scan of GPR image, extracts hyperbola features to remove the noise in B-scan data and predicts dielectric to determine the depth of the objects, 3) 3D GPR migration module which synchronizes the pose information with GPR scan data processed by DepthNet to reconstruct and visualize the 3D underground targets. Our proposed DepthNet processes the GPR data by removing the noise in B-scan image as well as predicting depth of subsurface objects. For DepthNet model training and testing, we collect the real GPR data in the concrete test pit at Geophysical Survey System Inc. (GSSI) and create the synthetic GPR data by using gprMax3.0 simulator. The dataset we create includes 350 labeled GPR images. The DepthNet achieves an average accuracy of 92.64% for B-scan feature detection and an 0.112 average error for underground target depth prediction. In addition, the experimental results verify that our proposed method improve the migration accuracy and performance in generating 3D GPR image compared with the traditional migration methods.
Jinglun Feng, Haiyan Wang 0019, Yifeng Song, Jizhong Xiao
ICRA3
2019 Towards Weakly Supervised Semantic Segmentation in 3D Graph-Structured Point Clouds of Wild Scenes
Haiyan Wang 0019, Xuejian Rong, Shuihua Wang, Yingli Tian
BMVC1
2019 Towards Accurate Instance-Level Text Spotting with Guided Attention
abstract
We tackle the text detection problem from the instance-aware segmentation perspective, in which text bounding boxes are directly extracted from segmentation results without location regression. Specifically, a text-specific attention model and a global enhancement block are introduced to enrich the semantics of text detection features. The attention model is trained with a weakly segmentation supervision signal and enforces the detector to focus on the text regions, while also suppressing the influence of neighboring background clutters. In conjunction with the attention model, a global enhancement block (GEB) is adapted to reason the relationship among different channels with channel-wise weights calibration. Our method achieves comparable performance with the recent state-of-the-arts on ICDAR2013, ICDAR2015, and ICDAR2017-MLT benchmark datasets.
Haiyan Wang 0019, Xuejian Rong, Yingli Tian
ICME1