EDBT 2026 Demo / reviewers in the wild / expert
Bo Yang 0027
dblp:46/999-27
· DBLP profile ↗
37ranked-venue papers
4as first author
29since 2021 · last 2026
0000-0002-2419-4140ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 34 · 3 first-author · 27 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 10 since 2021Systems, architecture and hardware · 5 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Computer networks · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Single-Stage fMRI-to-3D Reconstruction via Viewpoint-Aware Embedding and Hierarchical GuidanceabstractUnderstanding the neural basis of three-dimensional (3D) perception is a fundamental objective in cognitive neuroscience. Despite advances in decoding 2D visual stimuli from neural data, reconstructing high-fidelity 3D objects with detailed texture and geometry remains largely unexplored. In this work, we introduce NeuroSculptor3D, the first single-stage, end-to-end framework for reconstructing textured 3D shapes directly from brain activity. NeuroSculptor3D integrates a viewpoint-aware brain embedding module that captures fine-grained spatial variations across visual perspectives, and a hierarchical guidance mechanism that aligns brain-derived features with perceptual, semantic, and structural priors. Together, these components facilitate the generation of consistent multi-view embeddings, which are then decoded via TRELLIS to produce high-quality textured 3D reconstructions. Experiments on the fMRI-Shape dataset demonstrate that NeuroSculptor3D outperforms existing baselines across multiple settings, achieving significant improvements in both structural accuracy and semantic consistency. Code will be released to facilitate further research. Weihao Xia 0001, Bo Yang 0027, Alessandro Bozzon, Pan Wang 0005 |
AAAI | 4 |
| 2026 | GrowSP++: Growing Superpoints and Primitives for Unsupervised 3D Semantic SegmentationabstractWe study the problem of 3D semantic segmentation from raw point clouds. Unlike existing methods which primarily rely on a large amount of human annotations for training neural networks, we proposes GrowSP++, an unsupervised method to successfully identify complex semantic classes for every point in 3D scenes, without needing any type of human labels. Our method is composed of three major components: 1) a feature extractor incorporating 2D-3D feature distillation, 2) a superpoint constructor featuring progressively growing superpoints, and 3) a semantic primitive constructor with an additional growing strategy. The key to our method is the superpoint constructor together with the progressive growing strategy on both superpoints and semantic primitives, driving the feature extractor to progressively learn similar features for 3D points belonging to the same semantic class. We extensively evaluate our method on five challenging indoor and outdoor datasets, demonstrating state-of-the-art performance over all unsupervised baselines. We hope our work could inspire more advanced methods for unsupervised 3D semantic learning. Weisheng Dai, Bing Wang 0013, Bo Li 0037, Bo Yang 0027 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2025 | FreeGave: 3D Physics Learning from Dynamic Videos by Gaussian VelocityabstractIn this paper, we aim to model 3D scene geometry, appearance, and the underlying physics purely from multi-view videos. By applying various governing PDEs as PINN losses or incorporating physics simulation into neural networks, existing works often fail to learn complex physical motions at boundaries or require object priors such as masks or types. In this paper, we propose FreeGave to learn physics of complex dynamic 3D scenes without needing any object priors. The key to our approach is to introduce a physics code followed by a carefully designed divergence-free module for estimating a per-Gaussian velocity field, without relying on the inefficient PINN losses. Extensive experiments on three public datasets and a newly collected challenging real-world dataset demonstrate the superior performance of our method for future frame extrapolation and motion segmentation. Most notably, our investigation into the learned physics codes reveals that they truly learn meaningful 3D physical motion patterns in the absence of any human labels in training. Our code and data are available at https://github.com/vLAR-group/FreeGave Bo Yang 0027 |
CVPR | 4 |
| 2025 | LogoSP: Local-global Grouping of Superpoints for Unsupervised Semantic Segmentation of 3D Point CloudsabstractWe study the problem of unsupervised 3D semantic segmentation on raw point clouds without needing human labels in training. Existing methods usually formulate this problem into learning per-point local features followed by a simple grouping strategy, lacking the ability to discover additional and possibly richer semantic priors beyond local features. In this paper, we introduce LogoSP to learn 3D semantics from both local and global point features. The key to our approach is to discover 3D semantic information by grouping superpoints according to their global patterns in the frequency domain, thus generating highly accurate semantic pseudo-labels for training a segmentation network. Extensive experiments on two indoor and an outdoor datasets show that our LogoSP surpasses all existing unsupervised methods by large margins, achieving the state-of-the-art performance for unsupervised 3D semantic segmentation. Notably, our investigation into the learned global patterns reveals that they truly represent meaningful 3D semantics in the absence of human labels during training. Our code and data are available at https://github.com/vLAR-group/LogoSP Weisheng Dai, Hongtao Wen 0001, Bo Yang 0027 |
CVPR | 4 |
| 2025 | TRACE: Learning 3D Gaussian Physical Dynamics from Multi-View Videos
Bo Yang 0027 |
ICCV | 3 |
| 2025 | RayletDF: Raylet Distance Fields for Generalizable 3D Surface Reconstruction from Point Clouds or Gaussians
Shenxing Wei, Yafei Yang, Bo Yang 0027 |
ICCV | 5 |
| 2025 | GrabS: Generative Embodied Agent for 3D Object Segmentation without Scene SupervisionabstractWe study the hard problem of 3D object segmentation in complex point clouds
without requiring human labels of 3D scenes for supervision. By relying on the
similarity of pretrained 2D features or external signals such as motion to group 3D
points as objects, existing unsupervised methods are usually limited to identifying
simple objects like cars or their segmented objects are often inferior due to the
lack of objectness in pretrained features. In this paper, we propose a new two-
stage pipeline called GrabS. The core concept of our method is to learn generative
and discriminative object-centric priors as a foundation from object datasets in the
first stage, and then design an embodied agent to learn to discover multiple ob-
jects by querying against the pretrained generative priors in the second stage. We
extensively evaluate our method on two real-world datasets and a newly created
synthetic dataset, demonstrating remarkable segmentation performance, clearly
surpassing all existing unsupervised methods. Yafei Yang, Hongtao Wen 0001, Bo Yang 0027 |
ICLR | 4 |
| 2025 | unMORE: Unsupervised Multi-Object Segmentation via Center-Boundary ReasoningabstractWe study the challenging problem of unsupervised multi-object segmentation on single images. Existing methods, which rely on image reconstruction objectives to learn objectness or leverage pretrained image features to group similar pixels, often succeed only in segmenting simple synthetic objects or discovering a limited number of real-world objects. In this paper, we introduce unMORE, a novel two-stage pipeline designed to identify many complex objects in real-world images. The key to our approach involves explicitly learning three levels of carefully defined object-centric representations in the first stage. Subsequently, our multi-object reasoning module utilizes these learned object priors to discover multiple objects in the second stage. Notably, this reasoning module is entirely network-free and does not require human labels. Extensive experiments demonstrate that unMORE significantly outperforms all existing unsupervised methods across 6 real-world benchmark datasets, including the challenging COCO dataset, achieving state-of-the-art object segmentation results. Remarkably, our method excels in crowded images where all baselines collapse. Our code and data are available at https://github.com/vLAR-group/unMORE. Yafei Yang, Bo Yang 0027 |
ICML | 3 |
| 2024 | OSN: Infinite Representations of Dynamic 3D Scenes from Monocular VideosabstractIt has long been challenging to recover the underlying dynamic 3D scene representations from a monocular RGB video. Existing works formulate this problem into finding a single most plausible solution by adding various constraints such as depth priors and strong geometry constraints, ignoring the fact that there could be infinitely many 3D scene representations corresponding to a single dynamic video. In this paper, we aim to learn all plausible 3D scene configurations that match the input video, instead of just inferring a specific one. To achieve this ambitious goal, we introduce a new framework, called OSN. The key to our approach is a simple yet innovative object scale network together with a joint optimization module to learn an accurate scale range for every dynamic 3D object. This allows us to sample as many faithful 3D scene configurations as possible. Extensive experiments show that our method surpasses all baselines and achieves superior accuracy in dynamic novel view synthesis on multiple synthetic and real-world datasets. Most notably, our method demonstrates a clear advantage in learning fine-grained 3D scene geometry. Bo Yang 0027 |
ICML | 3 |
| 2024 | Learning to Catch Reactive Objects with a Behavior PredictorabstractTracking and catching moving objects is an important ability for robots in a dynamic world. Whilst some objects have highly predictable state evolution e.g., the ballistic trajectory of a tennis ball, reactive targets alter their behavior in response to motion of the manipulator. Reactive applications range from gently capturing living animals such as snakes or fish for biological investigations, to smoothly interacting with and assisting a person. Existing works for dynamic catching usually perform target prediction followed by planning, but seldom account for highly non-linear reactive behaviors. Alternatively, Reinforcement Learning (RL) based methods simply treat the target and its motion as part of the observation of the world-state, but perform poorly due to the weak reward signal. In this work, we blend the approach of an explicit, yet learned, target state predictor with RL. We further show how a tightly coupled predictor which ‘observes’ the state of the robot leads to significantly improved anticipatory action, especially with targets that seek to evade the robot following a simple policy. Experiments show that our method achieves an 86.4% (open plane area) and a 73.8% (room) success rate on evasive objects, outperforming monolithic reinforcement learning and other techniques. We also demonstrate the efficacy of our approach across varied targets and trajectories. All code, data, and additional videos are at this GitHub link: https://kl-research.github.io/dyncatch. Kai Lu 0003, Jia-Xing Zhong, Bo Yang 0027, Bing Wang 0013, Andrew Markham |
ICRA | 3 |
| 2024 | Benchmarking and Analysis of Unsupervised Object Segmentation from Real-World Single ImagesabstractAbstract In this paper, we study the problem of unsupervised object segmentation from single images. We do not introduce a new algorithm, but systematically investigate the effectiveness of existing unsupervised models on challenging real-world images. We first introduce seven complexity factors to quantitatively measure the distributions of background and foreground object biases in appearance and geometry for datasets with human annotations. With the aid of these factors, we empirically find that, not surprisingly, existing unsupervised models fail to segment generic objects in real-world images, although they can easily achieve excellent performance on numerous simple synthetic datasets, due to the vast gap in objectness biases between synthetic and real images. By conducting extensive experiments on multiple groups of ablated real-world datasets, we ultimately find that the key factors underlying the failure of existing unsupervised models on real-world images are the challenging distributions of background and foreground object biases in appearance and geometry. Because of this, the inductive biases introduced in existing unsupervised models can hardly capture the diverse object distributions. Our research results suggest that future work should exploit more explicit objectness biases in the network design. Yafei Yang, Bo Yang 0027 |
Int. J. Comput. Vis. | 2 |
| 2024 | Unsupervised 3D Object Segmentation of Point Clouds by Geometry ConsistencyabstractIn this paper, we study the problem of 3D object segmentation from raw point clouds. Unlike existing methods which usually require a large amount of human annotations for full supervision, we propose the first unsupervised method, called OGC, to simultaneously identify multiple 3D objects in a single forward pass, without needing any type of human annotations. The key to our approach is to fully leverage the dynamic motion patterns over sequential point clouds as supervision signals to automatically discover rigid objects. Our method consists of three major components, 1) the object segmentation network to directly estimate multi-object masks from a single point cloud frame, 2) the auxiliary self-supervised scene flow estimator, and 3) our core object geometry consistency component. By carefully designing a series of loss functions, we effectively take into account the multi-object rigid consistency and the object shape invariance in both temporal and spatial scales. This allows our method to truly discover the object geometry even in the absence of annotations. We extensively evaluate our method on five datasets, demonstrating the superior performance for object part instance segmentation and general object segmentation in both indoor and the challenging outdoor scenarios. Bo Yang 0027 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2024 | A Framework for Graphical GNSS Multipath and NLOS MitigationabstractPositioning in urban areas is still a challenge due to non-line-of-sight (NLOS) and multipath reception. This paper explores the geometrical characteristics of the GNSS ranging measurement by a graphical representation to better indicate the pseudorange consistency, which can be used to mitigate the NLOS and multipath receptions. The graphical representation is created by the grid-based method combined with the single differenced technique, which is called the single differenced residual map (SDRes Map). With the graphical properties of the SDRes Map, four main focuses of the NLOS/multipath problems, including positioning, signal status prediction, satellite weighting calculation, and NLOS/multipath error calculation, are able to be tackled simultaneously and demonstrated to have superior performance against the conventional or even state-of-the-art method methods. Penghui Xu, Guohao Zhang, Yihan Zhong, Bo Yang 0027, Li-Ta Hsu |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2023 | GrowSP: Unsupervised Semantic Segmentation of 3D Point CloudsabstractWe study the problem of 3D semantic segmentation from raw point clouds. Unlike existing methods which primarily rely on a large amount of human annotations for training neural networks, we propose the first purely unsupervised method, called GrowSP, to successfully identify complex semantic classes for every point in 3D scenes, without needing any type of human labels or pretrained models. The key to our approach is to discover 3D semantic elements via progressive growing of superpoints. Our method consists of three major components, 1) the feature extractor to learn per-point features from input point clouds, 2) the superpoint constructor to progressively grow the sizes of superpoints, and 3) the semantic primitive clustering module to group superpoints into semantic elements for the final semantic segmentation. We extensively evaluate our method on multiple datasets, demonstrating superior performance over all unsupervised baselines and approaching the classic fully-supervised PointNet. We hope our work could inspire more advanced methods for unsupervised 3D semantic learning. Bo Yang 0027, Bing Wang 0013, Bo Li 0037 |
CVPR | 2 |
| 2023 | DM-NeRF: 3D Scene Geometry Decomposition and Manipulation from 2D Images
Bing Wang 0013, Bo Yang 0027 |
ICLR | 3 |
| 2023 | Decoupling Skill Learning from Robotic Control for Generalizable Object ManipulationabstractRecent works in robotic manipulation through reinforcement learning (RL) or imitation learning (IL) have shown potential for tackling a range of tasks e.g., opening a drawer or a cupboard. However, these techniques generalize poorly to unseen objects. We conjecture that this is due to the high-dimensional action space for joint control. In this paper, we take an alternative approach and separate the task of learning ‘what to do’ from ‘how to do it’ i.e., whole-body control. We pose the RL problem as one of determining the skill dynamics for a disembodied virtual manipulator interacting with articulated objects. The whole-body robotic kinematic control is optimized to execute the high-dimensional joint motion to reach the goals in the workspace. It does so by solving a quadratic programming (QP) model with robotic singularity and kinematic constraints. Our experiments on manipulating complex articulated objects show that the proposed approach is more generalizable to unseen objects with large intra-class variations, outperforming previous approaches. The evaluation results indicate that our approach generates more compliant robotic motion and outperforms the pure RL and IL baselines in task success rates. Additional information and videos are available at https://kl-research.github.io/decoupskill. Kai Lu 0003, Bo Yang 0027, Bing Wang 0013, Andrew Markham |
ICRA | 2 |
| 2023 | NVFi: Neural Velocity Fields for 3D Physics Learning from Dynamic VideosabstractIn this paper, we aim to model 3D scene dynamics from multi-view videos. Unlike the majority of existing works which usually focus on the common task of novel view synthesis within the training time period, we propose to simultaneously learn the geometry, appearance, and physical velocity of 3D scenes only from video frames, such that multiple desirable applications can be supported, including future frame extrapolation, unsupervised 3D semantic scene decomposition, and dynamic motion transfer. Our method consists of three major components, 1) the keyframe dynamic radiance field, 2) the interframe velocity field, and 3) a joint keyframe and interframe optimization module which is the core of our framework to effectively train both networks. To validate our method, we further introduce two dynamic 3D datasets: 1) Dynamic Object dataset, and 2) Dynamic Indoor Scene dataset. We conduct extensive experiments on multiple datasets, demonstrating the superior performance of our method over all baselines, particularly in the critical tasks of future frame extrapolation and unsupervised 3D semantic scene decomposition. Bo Yang 0027 |
NeurIPS | 3 |
| 2023 | RayDF: Neural Ray-surface Distance Fields with Multi-view ConsistencyabstractIn this paper, we study the problem of continuous 3D shape representations. The majority of existing successful methods are coordinate-based implicit neural representations. However, they are inefficient to render novel views or recover explicit surface points. A few works start to formulate 3D shapes as ray-based neural functions, but the learned structures are inferior due to the lack of multi-view geometry consistency. To tackle these challenges, we propose a new framework called RayDF. It consists of three major components: 1) the simple ray-surface distance field, 2) the novel dual-ray visibility classifier, and 3) a multi-view consistency optimization module to drive the learned ray-surface distances to be multi-view geometry consistent. We extensively evaluate our method on three public datasets, demonstrating remarkable performance in 3D surface point reconstruction on both synthetic and challenging real-world 3D scenes, clearly surpassing existing coordinate-based and ray-based baselines. Most notably, our method achieves a 1000x faster speed than coordinate-based methods to render an 800x800 depth image, showing the superiority of our method for 3D shape representation. Our code and data are available at https://github.com/vLAR-group/RayDF Zhuoman Liu, Bo Yang 0027, Yan Luximon |
NeurIPS | 2 |
| 2023 | You Only Train Once: Learning General and Distinctive 3D Local DescriptorsabstractExtracting distinctive, robust, and general 3D local features is essential to downstream tasks such as point cloud registration. However, existing methods either rely on noise-sensitive handcrafted features, or depend on rotation-variant neural architectures. It remains challenging to learn robust and general local feature descriptors for surface matching. In this paper, we propose a new, simple yet effective neural network, termed SpinNet, to extract local surface descriptors which are rotation-invariant whilst sufficiently distinctive and general. A Spatial Point Transformer is first introduced to embed the input local surface into an elaborate cylindrical representation (SO(2) rotation-equivariant), further enabling end-to-end optimization of the entire framework. A Neural Feature Extractor, composed of point-based and 3D cylindrical convolutional layers, is then presented to learn representative and general geometric patterns. An invariant layer is finally used to generate rotation-invariant feature descriptors. Extensive experiments on both indoor and outdoor datasets demonstrate that SpinNet outperforms existing state-of-the-art techniques by a large margin. More critically, it has the best generalization ability across unseen scenarios with different sensor modalities. Sheng Ao, Yulan Guo, Qingyong Hu, Bo Yang 0027, Andrew Markham, Zengping Chen |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2023 | SemanticFlow: Semantic Segmentation of Sequential LiDAR Point Clouds From Sparse Frame AnnotationsabstractSequential point clouds acquired by light detection and ranging (LiDAR) technology provide accurate spatial information for environmental sensing. However, semantic segmentation of point cloud sequences relies on many manual point-wise annotations, which are error-prone and expensive. Existing mainstream weakly supervised methods tackle this by reducing the percentage of labeled points, but they are mostly designed for static indoor scenes and are hard to apply practically. From the viewpoint of realistic annotation procedures and the nature of point cloud sequences, this paper proposes a novel semantic segmentation method, SemanticFlow, for LiDAR point cloud sequences using sparse frames with annotations. The proposed method achieves competitive performance compared with fully supervised methods. Specifically, we designed a bidirectional cross-frame pseudo label propagation module that uses scene flow to learn the correlation and propagate pseudo labels across neighboring frames. In addition, a label refinement mechanism is proposed to select reliable pseudo labels for learning. Extensive experiments on SemanticKITTI, SemanticPOSS, and Synthia 4D datasets demonstrate that our sparse frame annotation method is compatible with some fully supervised counterparts. Junhao Zhao, Chenglu Wen, Bo Yang 0027, Yulan Guo, Cheng Wang 0003 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2022 | SQN: Weakly-Supervised Semantic Segmentation of Large-Scale 3D Point Clouds
Qingyong Hu, Bo Yang 0027, Guangchi Fang, Yulan Guo, Ales Leonardis, Agathoniki Trigoni, Andrew Markham |
ECCV (27) | 2 |
| 2022 | OGC: Unsupervised 3D Object Segmentation from Rigid Dynamics of Point CloudsabstractIn this paper, we study the problem of 3D object segmentation from raw point clouds. Unlike all existing methods which usually require a large amount of human annotations for full supervision, we propose the first unsupervised method, called OGC, to simultaneously identify multiple 3D objects in a single forward pass, without needing any type of human annotations. The key to our approach is to fully leverage the dynamic motion patterns over sequential point clouds as supervision signals to automatically discover rigid objects. Our method consists of three major components, 1) the object segmentation network to directly estimate multi-object masks from a single point cloud frame, 2) the auxiliary self-supervised scene flow estimator, and 3) our core object geometry consistency component. By carefully designing a series of loss functions, we effectively take into account the multi-object rigid consistency and the object shape invariance in both temporal and spatial scales. This allows our method to truly discover the object geometry even in the absence of annotations. We extensively evaluate our method on five datasets, demonstrating the superior performance for object part instance segmentation and general object segmentation in both indoor and the challenging outdoor scenarios. Bo Yang 0027 |
NeurIPS | 2 |
| 2022 | Promising or Elusive? Unsupervised Object Segmentation from Real-world Single ImagesabstractIn this paper, we study the problem of unsupervised object segmentation from single images. We do not introduce a new algorithm, but systematically investigate the effectiveness of existing unsupervised models on challenging real-world images. We firstly introduce four complexity factors to quantitatively measure the distributions of object- and scene-level biases in appearance and geometry for datasets with human annotations. With the aid of these factors, we empirically find that, not surprisingly, existing unsupervised models catastrophically fail to segment generic objects in real-world images, although they can easily achieve excellent performance on numerous simple synthetic datasets, due to the vast gap in objectness biases between synthetic and real images. By conducting extensive experiments on multiple groups of ablated real-world datasets, we ultimately find that the key factors underlying the colossal failure of existing unsupervised models on real-world images are the challenging distributions of object- and scene-level biases in appearance and geometry. Because of this, the inductive biases introduced in existing unsupervised models can hardly capture the diverse object distributions. Our research results suggest that future work should exploit more explicit objectness biases in the network design. Yafei Yang, Bo Yang 0027 |
NeurIPS | 2 |
| 2022 | SensatUrban: Learning Semantics from Urban-Scale Photogrammetric Point Clouds
Qingyong Hu, Bo Yang 0027, Sheikh Khalid, Agathoniki Trigoni, Andrew Markham |
Int. J. Comput. Vis. | 2 |
| 2022 | Learning Semantic Segmentation of Large-Scale Point Clouds With Random SamplingabstractWe study the problem of efficient semantic segmentation of large-scale 3D point clouds. By relying on expensive sampling techniques or computationally heavy pre/post-processing steps, most existing approaches are only able to be trained and operate over small-scale point clouds. In this paper, we introduce RandLA-Net, an efficient and lightweight neural architecture to directly infer per-point semantics for large-scale point clouds. The key to our approach is to use random point sampling instead of more complex point selection approaches. Although remarkably computation and memory efficient, random sampling can discard key features by chance. To overcome this, we introduce a novel local feature aggregation module to progressively increase the receptive field for each 3D point, thereby effectively preserving geometric details. Comparative experiments show that our RandLA-Net can process 1 million points in a single pass up to 200× faster than existing approaches. Moreover, extensive experiments on five large-scale point cloud datasets, including Semantic3D, SemanticKITTI, Toronto3D, NPM3D and S3DIS, demonstrate the state-of-the-art semantic segmentation performance of our RandLA-Net. Qingyong Hu, Bo Yang 0027, Linhai Xie, Stefano Rosa, Yulan Guo, Zhihua Wang 0005, Agathoniki Trigoni, Andrew Markham |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2021 | SpinNet: Learning a General Surface Descriptor for 3D Point Cloud RegistrationabstractExtracting robust and general 3D local features is key to downstream tasks such as point cloud registration and reconstruction. Existing learning-based local descriptors are either sensitive to rotation transformations, or rely on classical handcrafted features which are neither general nor representative. In this paper, we introduce a new, yet conceptually simple, neural architecture, termed SpinNet, to extract local features which are rotationally invariant whilst sufficiently informative to enable accurate registration. A Spatial Point Transformer is first introduced to map the input local surface into a carefully designed cylindrical space, enabling end-to-end optimization with SO(2) equivariant representation. A Neural Feature Extractor which leverages the powerful point-based and 3D cylindrical convolutional neural layers is then utilized to derive a compact and representative descriptor for matching. Extensive experiments on both indoor and outdoor datasets demonstrate that SpinNet outperforms existing state-of-the-art techniques by a large margin. More critically, it has the best generalization ability across unseen scenarios with different sensor modalities. The code is available at https://github.com/QingyongHu/SpinNet. Sheng Ao, Qingyong Hu, Bo Yang 0027, Andrew Markham, Yulan Guo |
CVPR | 3 |
| 2021 | Towards Semantic Segmentation of Urban-Scale 3D Point Clouds: A Dataset, Benchmarks and ChallengesabstractAn essential prerequisite for unleashing the potential of supervised deep learning algorithms in the area of 3D scene understanding is the availability of large-scale and richly annotated datasets. However, publicly available datasets are either in relative small spatial scales or have limited semantic annotations due to the expensive cost of data acquisition and data annotation, which severely limits the development of fine-grained semantic understanding in the context of 3D point clouds. In this paper, we present an urban-scale photogrammetric point cloud dataset with nearly three billion richly annotated points, which is three times the number of labeled points than the existing largest photogrammetric point cloud dataset. Our dataset consists of large areas from three UK cities, covering about 7.6 km2of the city landscape. In the dataset, each 3D point is labeled as one of 13 semantic classes. We extensively evaluate the performance of state-of-the-art algorithms on our dataset and provide a comprehensive analysis of the results. In particular, we identify several key challenges towards urban-scale point cloud understanding. The dataset is available at https://github.com/QingyongHu/SensatUrban. Qingyong Hu, Bo Yang 0027, Sheikh Khalid, Agathoniki Trigoni, Andrew Markham |
CVPR | 2 |
| 2021 | GRF: Learning a General Radiance Field for 3D Representation and RenderingabstractWe present a simple yet powerful neural network that implicitly represents and renders 3D objects and scenes only from 2D observations. The network models 3D geometries as a general radiance field, which takes a set of 2D images with camera poses and intrinsics as input, constructs an internal representation for each point of the 3D space, and then renders the corresponding appearance and geometry of that point viewed from an arbitrary position. The key to our approach is to learn local features for each pixel in 2D images and to then project these features to 3D points, thus yielding general and rich point representations. We additionally integrate an attention mechanism to aggregate pixel features from multiple 2D views, such that visual occlusions are implicitly taken into account. Extensive experiments demonstrate that our method can generate high-quality and realistic novel views for novel objects, unseen categories and challenging real-world scenes. Alex Trevithick, Bo Yang 0027 |
ICCV | 2 |
| 2021 | RadarLoc: Learning to Relocalize in FMCW RadarabstractRelocalization is a fundamental task in the field of robotics and computer vision. There is considerable work in the field of deep camera relocalization, which directly estimates poses from raw images. However, learning-based methods have not yet been applied to the radar sensory data. In this work, we investigate how to exploit deep learning to predict global poses from Emerging Frequency-Modulated Continuous Wave (FMCW) radar scans. Specifically, we propose a novel end-to-end neural network with self-attention, termed RadarLoc, which is able to estimate 6-DoF global poses directly. We also propose to improve the localization performance by utilizing geometric constraints between radar scans. We validate our approach on the recently released challenging outdoor dataset Oxford Radar RobotCar. Comprehensive experiments demonstrate that the proposed method outperforms radar-based localization and deep camera relocalization methods by a significant margin. Wei Wang 0226, Pedro Porto Buarque de Gusmão, Bo Yang 0027, Andrew Markham, Agathoniki Trigoni |
ICRA | 3 |
| 2020 | RandLA-Net: Efficient Semantic Segmentation of Large-Scale Point CloudsabstractWe study the problem of efficient semantic segmentation for large-scale 3D point clouds. By relying on expensive sampling techniques or computationally heavy pre/post-processing steps, most existing approaches are only able to be trained and operate over small-scale point clouds. In this paper, we introduce RandLA-Net, an efficient and lightweight neural architecture to directly infer per-point semantics for large-scale point clouds. The key to our approach is to use random point sampling instead of more complex point selection approaches. Although remarkably computation and memory efficient, random sampling can discard key features by chance. To overcome this, we introduce a novel local feature aggregation module to progressively increase the receptive field for each 3D point, thereby effectively preserving geometric details. Extensive experiments show that our RandLA-Net can process 1 million points in a single pass with up to 200x faster than existing approaches. Moreover, our RandLA-Net clearly surpasses state-of-the-art approaches for semantic segmentation on two large-scale benchmarks Semantic3D and SemanticKITTI. Qingyong Hu, Bo Yang 0027, Linhai Xie, Stefano Rosa, Yulan Guo, Zhihua Wang 0005, Agathoniki Trigoni, Andrew Markham |
CVPR | 2 |
| 2020 | Robust Attentional Aggregation of Deep Feature Sets for Multi-view 3D ReconstructionabstractWe study the problem of recovering an underlying 3D shape from a set of images. Existing learning based approaches usually resort to recurrent neural nets, e.g., GRU, or intuitive pooling operations, e.g., max/mean poolings, to fuse multiple deep features encoded from input images. However, GRU based approaches are unable to consistently estimate 3D shapes given different permutations of the same set of input images as the recurrent unit is permutation variant. It is also unlikely to refine the 3D shape given more images due to the long-term memory loss of GRU. Commonly used pooling approaches are limited to capturing partial information, e.g., max/mean values, ignoring other valuable features. In this paper, we present a new feed-forward neural module, named AttSets , together with a dedicated training algorithm, named FASet , to attentively aggregate an arbitrarily sized deep feature set for multi-view 3D reconstruction. The AttSets module is permutation invariant, computationally efficient and flexible to implement, while the FASet algorithm enables the AttSets based network to be remarkably robust and generalize to an arbitrary number of input images. We thoroughly evaluate FASet and the properties of AttSets on multiple large public datasets. Extensive experiments show that AttSets together with FASet algorithm significantly outperforms existing aggregation approaches. Bo Yang 0027, Sen Wang 0002, Andrew Markham, Agathoniki Trigoni |
Int. J. Comput. Vis. | 1 |
| 2019 | DeepPCO: End-to-End Point Cloud Odometry through Deep Parallel Neural NetworkabstractOdometry is of key importance for localization in the absence of a map. There is considerable work in the area of visual odometry (VO), and recent advances in deep learning have brought novel approaches to VO, which directly learn salient features from raw images. These learning-based approaches have led to more accurate and robust VO systems. However, they have not been well applied to point cloud data yet. In this work, we investigate how to exploit deep learning to estimate point cloud odometry (PCO), which may serve as a critical component in point cloud-based downstream tasks or learning-based systems. Specifically, we propose a novel end-to-end deep parallel neural network called DeepPCO, which can estimate the 6-DOF poses using consecutive point clouds. It consists of two parallel sub-networks to estimate 3D translation and orientation respectively rather than a single neural network. We validate our approach on KITTI Visual Odometry/SLAM benchmark dataset with different baselines. Experiments demonstrate that the proposed approach achieves good performance in terms of pose accuracy. Wei Wang 0226, Muhamad Risqi Utama Saputra, Peijun Zhao, Pedro Porto Buarque de Gusmão, Bo Yang 0027, Changhao Chen, Andrew Markham, Agathoniki Trigoni |
IROS | 5 |
| 2019 | Learning Object Bounding Boxes for 3D Instance Segmentation on Point CloudsabstractWe propose a novel, conceptually simple and general framework for instance segmentation on 3D point clouds. Our method, called 3D-BoNet, follows the simple design philosophy of per-point multilayer perceptrons (MLPs). The framework directly regresses 3D bounding boxes for all instances in a point cloud, while simultaneously predicting a point-level mask for each instance. It consists of a backbone network followed by two parallel network branches for 1) bounding box regression and 2) point mask prediction. 3D-BoNet is single-stage, anchor-free and end-to-end trainable. Moreover, it is remarkably computationally efficient as, unlike existing approaches, it does not require any post-processing steps such as non-maximum suppression, feature sampling, clustering or voting. Extensive experiments show that our approach surpasses existing work on both ScanNet and S3DIS datasets while being approximately 10x more computationally efficient. Comprehensive ablation studies demonstrate the effectiveness of our design. Bo Yang 0027, Ronald Clark, Qingyong Hu, Sen Wang 0002, Andrew Markham, Agathoniki Trigoni |
NeurIPS | 1 |
| 2019 | Dense 3D Object Reconstruction from a Single Depth ViewabstractIn this paper, we propose a novel approach, 3D-RecGAN++, which reconstructs the complete 3D structure of a given object from a single arbitrary depth view using generative adversarial networks. Unlike existing work which typically requires multiple views of the same object or class labels to recover the full 3D geometry, the proposed 3D-RecGAN++ only takes the voxel grid representation of a depth view of the object as input, and is able to generate the complete 3D occupancy grid with a high resolution of 2563by recovering the occluded/missing regions. The key idea is to combine the generative capabilities of 3D encoder-decoder and the conditional adversarial networks framework, to infer accurate and fine-grained 3D structures of objects in high-dimensional voxel space. Extensive experiments on large synthetic datasets and real-world Kinect datasets show that the proposed 3D-RecGAN++ significantly outperforms the state of the art in single view 3D object reconstruction, and is able to reconstruct unseen types of objects. Bo Yang 0027, Stefano Rosa, Andrew Markham, Agathoniki Trigoni, Hongkai Wen 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2018 | DEFO-NET: Learning Body Deformation Using Generative Adversarial NetworksabstractModelling the physical properties of everyday objects is a fundamental prerequisite for autonomous robots. We present a novel generative adversarial network (DEFO-NET), able to predict body deformations under external forces from a single RGB-D image. The network is based on an invertible conditional Generative Adversarial Network (IcGAN) and is trained on a collection of different objects of interest generated by a physical finite element model simulator. Defo-netinherits the generalisation properties of GANs. This means that the network is able to reconstruct the whole 3-D appearance of the object given a single depth view of the object and to generalise to unseen object configurations. Contrary to traditional finite element methods, our approach is fast enough to be used in real-time applications. We apply the network to the problem of safe and fast navigation of mobile robots carrying payloads over different obstacles and floor materials. Experimental results in real scenarios show how a robot equipped with an RGB-D camera can use the network to predict terrain deformations under different payload configurations and use this to avoid unsafe areas. Zhihua Wang 0005, Stefano Rosa, Linhai Xie, Bo Yang 0027, Sen Wang 0002, Agathoniki Trigoni, Andrew Markham |
ICRA | 4 |
| 2018 | 3D-PhysNet: Learning the Intuitive Physics of Non-Rigid Object DeformationsabstractThe ability to interact and understand the environment is a fundamental prerequisite for a wide range of applications from robotics to augmented reality. In particular, predicting how deformable objects will react to applied forces in real time is a significant challenge. This is further confounded by the fact that shape information about encountered objects in the real world is often impaired by occlusions, noise and missing regions e.g. a robot manipulating an object will only be able to observe a partial view of the entire solid. In this work we present a framework, 3D-PhysNet, which is able to predict how a three-dimensional solid will deform under an applied force using intuitive physics modelling. In particular, we propose a new method to encode the physical properties of the material and the applied force, enabling generalisation over materials. The key is to combine deep variational autoencoders with adversarial training, conditioned on the applied force and the material properties.We further propose a cascaded architecture that takes a single 2.5D depth view of the object and predicts its deformation. Training data is provided by a physics simulator. The network is fast enough to be used in real-time applications from partial views. Experimental results show the viability and the generalisation properties of the proposed architecture. Zhihua Wang 0005, Stefano Rosa, Bo Yang 0027, Sen Wang 0002, Agathoniki Trigoni, Andrew Markham |
IJCAI | 3 |
| 2016 | Updating Wireless Signal Map with Bayesian Compressive SensingabstractIn a wireless system, a signal map shows the signal strength at different locations termed reference points (RPs). As access points (APs) and their transmission power may change over time, keeping an updated signal map is important for applications such as Wi-Fi optimization and indoor localization. Traditionally, the signal map is obtained by a full site survey, which is time-consuming and costly. We address in this paper how to efficiently update a signal map given sparse samples randomly crowdsourced in the space (e.g., by signal monitors, explicit human input, or implicit user participation). Bo Yang 0027, Suining He, Shueng-Han Gary Chan |
MSWiM | 1 |