Maosheng Ye

dblp:126/6604 · DBLP profile ↗
← Back
15ranked-venue papers
5as first author
10since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 5 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 5 first-author · 9 since 2021Systems, architecture and hardware · 4 · 1 since 2021Human-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2025 End-to-End HOI Reconstruction Transformer with Graph-based Encoding
abstract
With the diversification of human-object interaction (HOI) applications and the success of capturing human meshes, HOI reconstruction has gained widespread attention. Existing mainstream HOI reconstruction methods often rely on explicitly modeling interactions between humans and objects. However, such a way leads to a natural conflict between 3D mesh reconstruction, which emphasizes global structure, and fine-grained contact reconstruction, which focuses on local details. To address the limitations of explicit modeling, we propose the End-to-End HOI Reconstruction Transformer with Graph-based Encoding (HOI-TG). It implicitly learns the interaction between humans and objects by leveraging self-attention mechanisms. Within the transformer architecture, we devise graph residual blocks to aggregate the topology among vertices of different spatial structures. This dual focus effectively balances global and local representations. Without bells and whistles, HOI-TG achieves state-of-the-art performance on BEHAVE and InterCap datasets. Particularly on the challenging InterCap dataset, our method improves the reconstruction results for human and object meshes by 8.9% and 8.6%, respectively.
Zhenrong Wang, Sihan Ma, Maosheng Ye, Yibing Zhan, Dongjiang Li
CVPR4
2025 Hints of Prompt: Enhancing Visual Representation for Multimodal LLMs in Autonomous Driving
abstract
In light of the dynamic nature of autonomous driving environments and stringent safety requirements, general MLLMs combined with CLIP alone often struggle to accurately represent driving-specific scenarios, particularly in complex interactions and long-tail cases. To address this, we propose the Hints of Prompt (HoP) framework, which introduces three key enhancements: Affinity hint to emphasize instance-level structure by strengthening token-wise connections, Semantic hint to incorporate high-level information relevant to driving-specific cases, such as complex interactions among vehicles and traffic signs, and Question hint to align visual features with the query context, focusing on question-relevant regions. These hints are fused through a Hint Fusion module, enriching visual representations by capturing driving-related representations with limited domain data, ensuring faster adaptation to driving scenarios. Extensive experiments confirm the effectiveness of the HoP framework, showing that it significantly outperforms previous state-of-the-art methods in all key metrics.
Zhanning Gao, Maosheng Ye, Qifeng Chen 0001, Tongyi Cao, Honggang Qi
ICCV4
2024 Learning High-Resolution Vector Representation from Multi-camera Images for 3D Object Detection
Shuangjie Xu, Maosheng Ye, Zian Qian, Xiaoyi Zou, Dit-Yan Yeung, Qifeng Chen 0001
ECCV (35)3
2024 PPAD: Iterative Interactions of Prediction and Planning for End-to-End Autonomous Driving
Maosheng Ye, Shuangjie Xu, Tongyi Cao, Qifeng Chen 0001
ECCV (35)2
2024 Cross-Cluster Shifting for Efficient and Effective 3D Object Detection in Autonomous Driving
abstract
We present a new 3D point-based detector model, named Shift-SSD, for precise 3D object detection in autonomous driving. Traditional point-based 3D object detectors often employ architectures that rely on a progressive downsampling of points. While this method effectively reduces computational demands and increases receptive fields, it will compromise the preservation of crucial non-local information for accurate 3D object detection, especially in the complex driving scenarios. To address this, we introduce an intriguing Cross-Cluster Shifting operation to unleash the representation capacity of the point-based detector by efficiently modeling longer-range inter-dependency while including only a negligible overhead. Concretely, the Cross-Cluster Shifting operation enhances the conventional design by shifting partial channels from neighboring clusters, which enables richer interaction with non-local regions and thus enlarges the receptive field of clusters. We conduct extensive experiments on the KITTI, Waymo, and nuScenes datasets, and the results demonstrate the state-of-the-art performance of Shift-SSD in both detection accuracy and runtime efficiency.
Kien T. Pham 0001, Maosheng Ye, Qifeng Chen 0001
ICRA3
2023 Bootstrap Motion Forecasting With Self-Consistent Constraints
abstract
We present a novel framework to bootstrap Motion forecastIng with Self-consistent Constraints (MISC). The motion forecasting task aims at predicting future trajectories of vehicles by incorporating spatial and temporal information from the past. A key design of MISC is the proposed Dual Consistency Constraints that regularize the predicted trajectories under spatial and temporal perturbation during training. Also, to model the multi-modality in motion forecasting, we design a novel self-ensembling scheme to obtain accurate teacher targets to enforce the self-constraints with multi-modality supervision. With explicit constraints from multiple teacher targets, we observe a clear improvement in the prediction performance. Extensive experiments on the Argoverse motion forecasting benchmark and Waymo Open Motion dataset show that MISC significantly outperforms the state-of-the-art methods. As the proposed strategies are general and can be easily incorporated into other motion forecasting approaches, we also demonstrate that our proposed scheme consistently improves the prediction performance of several existing methods.
Maosheng Ye, Jiamiao Xu, Xunnong Xu, Tengfei Wang 0002, Tongyi Cao, Qifeng Chen 0001
ICCV1
2022 Sparse Cross-Scale Attention Network for Efficient LiDAR Panoptic Segmentation
abstract
Two major challenges of 3D LiDAR Panoptic Segmentation (PS) are that point clouds of an object are surface-aggregated and thus hard to model the long-range dependency especially for large instances, and that objects are too close to separate each other. Recent literature addresses these problems by time-consuming grouping processes such as dual-clustering, mean-shift offsets and etc., or by bird-eye-view (BEV) dense centroid representation that downplays geometry. However, the long-range geometry relationship has not been sufficiently modeled by local feature learning from the above methods. To this end, we present SCAN, a novel sparse cross-scale attention network to first align multi-scale sparse features with global voxel-encoded attention to capture the long-range relationship of instance context, which is able to boost the regression accuracy of the over-segmented large objects. For the surface-aggregated points, SCAN adopts a novel sparse class-agnostic representation of instance centroids, which can not only maintain the sparsity of aligned features to solve the under-segmentation on small objects, but also reduce the computation amount of the network through sparse convolution. Our method outperforms previous methods by a large margin in the SemanticKITTI dataset for the challenging 3D PS task, achieving 1st place with a real-time inference speed.
Shuangjie Xu, Rui Wan, Maosheng Ye, Xiaoyi Zou, Tongyi Cao
AAAI3
2022 Efficient Point Cloud Segmentation with Geometry-Aware Sparse Networks
Maosheng Ye, Rui Wan, Shuangjie Xu, Tongyi Cao, Qifeng Chen 0001
ECCV (39)1
2021 TPCN: Temporal Point Cloud Networks for Motion Forecasting
abstract
We propose the Temporal Point Cloud Networks (TPCN), a novel and flexible framework with joint spatial and temporal learning for trajectory prediction. Unlike existing approaches that rasterize agents and map information as 2D images or operate in a graph representation, our approach extends ideas from point cloud learning with dynamic temporal learning to capture both spatial and temporal information by splitting trajectory prediction into both spatial and temporal dimensions. In the spatial dimension, agents can be viewed as an unordered point set, and thus it is straightforward to apply point cloud learning techniques to model agents’ locations. While the spatial dimension does not take kinematic and motion information into account, we further propose dynamic temporal learning to model agents’ motion over time. Experiments on the Argoverse motion forecasting benchmark show that our approach achieves state-of-the-art results.
Maosheng Ye, Tongyi Cao, Qifeng Chen 0001
CVPR1
2021 DRINet: A Dual-Representation Iterative Learning Network for Point Cloud Segmentation
abstract
We present a novel and flexible architecture for point cloud segmentation with dual-representation iterative learning. In point cloud processing, different representations have their own pros and cons. Thus, finding suitable ways to represent point cloud data structure while keeping its own internal physical property such as permutation and scale-invariant is a fundamental problem. Therefore, we propose our work, DRINet, which serves as the basic network structure for dual-representation learning with great flexibility at feature transferring and less computation cost, especially for large-scale point clouds. DRINet mainly consists of two modules called Sparse Point-Voxel Feature Extraction and Sparse Voxel-Point Feature Extraction. By utilizing these two modules iteratively, features can be propagated between two different representations. We further propose a novel multi-scale pooling layer for pointwise locality learning to improve context information propagation. Our network achieves state-of-the-art results for point cloud classification and segmentation tasks on several datasets while maintaining high runtime efficiency. For large-scale outdoor scenarios, our method outperforms state-of-the-art methods with a real-time inference time of 62ms per frame.
Maosheng Ye, Shuangjie Xu, Tongyi Cao, Qifeng Chen 0001
ICCV1
2020 HVNet: Hybrid Voxel Network for LiDAR Based 3D Object Detection
abstract
We present Hybrid Voxel Network (HVNet), a novel one-stage unified network for point cloud based 3D object detection for autonomous driving. Recent studies show that 2D voxelization with per voxel PointNet style feature extractor leads to accurate and efficient detector for large 3D scenes. Since the size of the feature map determines the computation and memory cost, the size of the voxel becomes a parameter that is hard to balance. A smaller voxel size gives a better performance, especially for small objects, but a longer inference time. A larger voxel can cover the same area with a smaller feature map, but fails to capture intricate features and accurate location for smaller objects. We present a Hybrid Voxel network that solves this problem by fusing voxel feature encoder (VFE) of different scales at point-wise level and project into multiple pseudo-image feature maps. We further propose an attentive voxel feature encoding that outperforms plain VFE and a feature fusion pyramid network to aggregate multi-scale information at feature map level. Experiments on the KITTI benchmark show that a single HVNet achieves the best mAP among all existing methods with a real time inference speed of 31Hz.
Maosheng Ye, Shuangjie Xu, Tongyi Cao
CVPR1
2020 Review of the calibration of a structured light system
abstract
The accuracy of a structured light system largely depends on its calibration precision. This paper makes a review of some representative calibration approaches in the literature. Based on the techniques used in the calibration process, these approaches are divided into four categories: the calibration methods based on well-designed calibration target, invariance of the cross ratio, pseudo camera and other methods. Then the four categories are introduced separately.
Jian Zhang 0055, Maosheng Ye, Hong Mi
IECON3
2020 Research on point-cloud collection and 3D model reconstruction
abstract
Point-cloud collection is used to collect 3D surface features from the object. 3D reconstruction can form a visual 3D model based on point-cloud. They are the important parts of 3D surface measurement. 3D surface measurement can effectively help enterprises shorten the design cycle, improve product quality, save labor costs, and improve the competitiveness of enterprises. Optical image measurement is a branch of 3D surface measurement. Because optical image measurement has the advantages of non-contact, high speed, high degree of automation and good flexibility, it has been researched and applied widely. Image processing and calibration technology are often used in optical image measurement. Image processing can extract valuable information from images, and calibration technology is necessary for mathematical model. Multi-line structured light has been widely used in the measurement.
Jianan Sheng, Jian Zhang 0055, Hong Mi, Maosheng Ye
IECON4
2020 GOSMatch: Graph-of-Semantics Matching for Detecting Loop Closures in 3D LiDAR data
abstract
Detecting loop closures in 3D Light Detection and Ranging (LiDAR) data is a challenging task since point-level methods always suffer from instability. This paper presents a semantic-level approach named GOSMatch to perform reliable place recognition. Our method leverages novel descriptors, which are generated from the spatial relationship between semantics, to perform frame description and data association. We also propose a coarse-to-fine strategy to efficiently search for loop closures. Besides, GOSMatch can give an accurate 6-DOF initial pose estimation once a loop closure is confirmed. Extensive experiments have been conducted on the KITTI odometry dataset and the results show that GOSMatch can achieve robust loop closure detection performance and outperform existing methods.
Yachen Zhu, Yanyang Ma, Long Chen 0005, Maosheng Ye, Lingxi Li 0001
IROS5
2008 Intelligent observer-based road surface condition detection and identification
abstract
Road surface condition is greatly dependent on the surface's friction coefficient. The abrupt change of the coefficient results in variation of wheel slip which likely leads to vehicle instability. Although the vehicle on-board sensors can measure the vehicle's velocities and yaw rate, the measurements, often containing noise and drift, are limited to the surface that the vehicle is engaged. In contrast, an effective observer can be used to estimate the vehicle dynamics for all possible surface conditions. This paper proposes a new observer, called Extended State Observer (ESO) to estimate the three quantities, and more importantly an additional quantity known as system dynamics. With the aid of the ESO, the following three tasks are performed: (1) noise filtering from the measurement data (2) detection and classification of surface condition change, and (3) identification of the road surface. Fuzzy logic was employed to quickly detect the change of road surface condition and further classify the surface; a neural network was employed to help determine the friction coefficient. The dynamic model used in this study can be applied to four-wheel independent drive vehicles. The presented methods were simulated when a vehicle encountered a significant change from a uniform-μ (i.e. uniform friction coefficient) surface to a split-μ surface (i.e. different friction coefficient on each side of the wheels) during cornering.
Paul P. Lin, Maosheng Ye, Kuo-Ming Lee
SMC2