VLDB 2026 Research / reviewers in the wild / expert
Yao Duan
dblp:147/6362
· DBLP profile ↗
7ranked-venue papers
3as first author
6since 2021 · last 2024
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 6 since 2021Artificial intelligence and machine learning · 2 · 2 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | DSGI-Net: Density-based Selective Grouping Point Cloud Learning Network for Indoor SceneabstractAbstract Indoor scene point clouds exhibit diverse distributions and varying levels of sparsity, characterized by more intricate geometry and occlusion compared to outdoor scenes or individual objects. Despite recent advancements in 3D point cloud analysis introducing various network architectures, there remains a lack of frameworks tailored to the unique attributes of indoor scenarios. To address this, we propose DSGI‐Net, a novel indoor scene point cloud learning network that can be integrated into existing models. The key innovation of this work is selectively grouping more informative neighbor points in sparse regions and promoting semantic consistency of the local area where different instances are in proximity but belong to distinct categories. Furthermore, our method encodes both semantic and spatial relationships between points in local regions to reduce the loss of local geometric details. Extensive experiments on the ScanNetv2, SUN RGB‐D, and S3DIS indoor scene benchmarks demonstrate that our method is straightforward yet effective. Xin Wen 0005, Yao Duan, Kai Xu 0004, Chenyang Zhu 0002 |
Comput. Graph. Forum | 2 |
| 2024 | THP: Tensor-field-driven hierarchical path planning for autonomous scene exploration with depth sensorsabstractIt is challenging to automatically explore an unknown 3D environment with a robot only equipped with depth sensors due to the limited field of view. We introduce THP, a tensor field-based framework for efficient environment exploration which can better utilize the encoded depth information through the geometric characteristics of tensor fields. Specifically, a corresponding tensor field is constructed incrementally and guides the robot to formulate optimal global exploration paths and a collision-free local movement strategy. Degenerate points generated during the exploration are adopted as anchors to formulate a hierarchical TSP for global path optimization. This novel strategy can help the robot avoid long-distance round trips more effectively while maintaining scanning completeness. Furthermore, the tensor field also enables a local movement strategy to avoid collision based on particle advection. As a result, the framework can eliminate massive, time-consuming recalculations of local movement paths. We have experimentally evaluate our method with a ground robot in 8 complex indoor scenes. Our method can on average achieve 14% better exploration efficiency and 21% better exploration completeness than state-of-the-art alternatives using LiDAR scans. Moreover, compared to similar methods, our method makes path decisions 39% faster due to our hierarchical exploration strategy. Yuefeng Xi, Chenyang Zhu 0002, Yao Duan, Renjiao Yi, Hongjun He, Kai Xu 0004 |
Comput. Vis. Media | 3 |
| 2023 | EFECL: Feature encoding enhancement with contrastive learning for indoor 3D object detectionabstractGood proposal initials are critical for 3D object detection applications. However, due to the significant geometry variation of indoor scenes, incomplete and noisy proposals are inevitable in most cases. Mining feature information among these “bad” proposals may mislead the detection. Contrastive learning provides a feasible way for representing proposals, which can align complete and incomplete/noisy proposals in feature space. The aligned feature space can help us build robust 3D representation even if bad proposals are given. Therefore, we devise a new contrast learning framework for indoor 3D object detection, called EFECL, that learns robust 3D representations by contrastive learning of proposals on two different levels. Specifically, we optimize both instance-level and category-level contrasts to align features by capturing instance-specific characteristics and semantic-aware common patterns. Furthermore, we propose an enhanced feature aggregation module to extract more general and informative features for contrastive learning. Evaluations on ScanNet V2 and SUN RGB-D benchmarks demonstrate the generalizability and effectiveness of our method, and our method can achieve 12.3% and 7.3% improvements on both datasets over the benchmark alternatives. The code and models are publicly available at https://github.com/YaraDuan/EFECL . Yao Duan, Renjiao Yi, Yuanming Gao, Kai Xu 0004, Chenyang Zhu 0002 |
Comput. Vis. Media | 1 |
| 2022 | DisARM: Displacement Aware Relation Module for 3D DetectionabstractWe introduce Displacement Aware Relation Module (DisARM), a novel neural network module for enhancing the performance of 3D object detection in point cloud scenes. The core idea is extracting the most principal contextual information is critical for detection while the target is incomplete or featureless. We find that relations between proposals provide a good representation to describe the context. However, adopting relations between all the object or patch proposals for detection is inefficient, and an imbalanced combination of local and global relations brings extra noise that could mislead the training. Rather than working with all relations, we find that training with relations only between the most representative ones, or an-chors, can significantly boost the detection performance. Good anchors should be semantic-aware with no ambiguity and able to describe the whole layout of a scene with no redundancy. To find the anchors, we first perform a preliminary relation anchor module with an objectness-aware sampling approach and then devise a displacement based module for weighing the relation importance for better utilization of contextual information. This lightweight relation module leads to significantly higher accuracy of object instance detection when being plugged into the state-of-the-art detectors. Evaluations on the public benchmarks of real-world scenes show that our method achieves the state-of-the-art performance on both SUN RGB-D and Scan-Net V2. The code and models are publicly available at https://github.com/YaraDuan/DisARM. Yao Duan, Chenyang Zhu 0002, Yuqing Lan, Renjiao Yi, Xinwang Liu 0002, Kai Xu 0004 |
CVPR | 1 |
| 2022 | ARM3D: Attention-based relation module for indoor 3D object detectionabstractRelation contexts have been proved to be useful for many challenging vision tasks. In the field of 3D object detection, previous methods have been taking the advantage of context encoding, graph embedding, or explicit relation reasoning to extract relation contexts. However, there exist inevitably redundant relation contexts due to noisy or low-quality proposals. In fact, invalid relation contexts usually indicate underlying scene misunderstanding and ambiguity, which may, on the contrary, reduce the performance in complex scenes. Inspired by recent attention mechanism like Transformer, we propose a novel 3D attention-based relation module (ARM3D). It encompasses object-aware relation reasoning to extract pair-wise relation contexts among qualified proposals and an attention module to distribute attention weights towards different relation contexts. In this way, ARM3D can take full advantage of the useful relation contexts and filter those less relevant or even confusing contexts, which mitigates the ambiguity in detection. We have evaluated the effectiveness of ARM3D by plugging it into several state-of-the-art 3D object detectors and showing more accurate and robust detection results. Extensive experiments show the capability and generalization of ARM3D on 3D object detection. Our source code is available at https://github.com/lanlan96/ARM3D . Yuqing Lan, Yao Duan, Chenyi Liu, Chenyang Zhu 0002, Yueshan Xiong, Hui Huang 0004, Kai Xu 0004 |
Comput. Vis. Media | 2 |
| 2021 | 3DRM: Pair-wise relation module for 3D object detection
Yuqing Lan, Yao Duan, Hui Huang 0004, Kai Xu 0004 |
Comput. Graph. | 2 |
| 2008 | Multi-objective design optimization of Surface Mount Permanent Magnet machine with particle swarm intelligenceabstractAn efficient multi-objective design method with Particle Swarm Optimization (PSO) is developed for Surface Mount Permanent Magnet machines to reduce the complexity in the PMSM machine design. First an analytical model of the PMSM machine’s geometry is developed and results are verified by finite element analysis. With proper design specification and assumption, the design input variables in this model can be reduced to as low as two, which significantly simplifies the optimization process. PSO is then applied to this analytical model. Compared to the traditional machine design methods, this proposed algorithm finds the optimized solution with fast computation and high convergence. Yao Duan, Ronald G. Harley, Thomas G. Habetler |
SIS | 1 |