Shuhui Bu

dblp:84/11232 · DBLP profile ↗
← Back
59ranked-venue papers
9as first author
15since 2021 · last 2026
0000-0002-9853-8313ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 26 · 4 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 24 · 5 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 12 · 6 since 2021Systems, architecture and hardware · 5 · 1 first-author · 2 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 CoMA-SLAM: Collaborative Multi-Agent Gaussian SLAM with Geometric Consistency
abstract
Although Gaussian scene representation has achieved remarkable success in tracking and mapping, most existing methods are confined to single-agent systems. Current multi-agent solutions typically rely on centralized architectures, which struggle to account for communication bandwidth constraints. Furthermore, the inherent depth ambiguity of 3D Gaussian splatting poses notable challenges in maintaining geometric consistency. To address these challenges, we introduce CoMA-SLAM, the first distributed multi-agent Gaussian SLAM framework. By leveraging 2D Gaussian surfels and robust initialization strategy, CoMA-SLAM enhances tracking accuracy and geometry consistency. It efficiently manages communication bandwidth while dynamically scaling with the number of agents. Through the integration of intra- and inter-loop closure, distributed keyframe optimization and submap centric update, our framework ensures global consistency and robustly alignment. Synthetic and real-world experiments demonstrate that CoMA-SLAM outperforms state-of-the-art methods in pose accuracy, rendering fidelity, and geometric consistency while maintaining competitive efficiency across distributed multi-agent systems. Notably, by avoiding data transmission to a centralized server, our method reduces communication bandwidth by 99.8% compared to centralized approaches.
Lin Chen 0042, Yongxin Su, Jvboxi Wang, Pengcheng Han, Zhenyu Xia, Shuhui Bu, Boni Hu, Shengqi Meng, Guangming Wang 0001
AAAI6
2025 CODE: COllaborative Visual-UWB SLAM for Online Large-Scale Metric DEnse Mapping
abstract
This paper presents a novel collaborative online dense mapping system for multiple Unmanned Aerial Vehicles (UAVs). The system confers two primary benefits: it facilitates simultaneous UAVs co-localization and real-time dense map reconstruction, and it recovers the metric scale even in GNSS-denied conditions. To achieve these advantages, Ultrawideband (UWB) measurements, monocular Visual Odometry (VO), and co-visibility observations are jointly employed to recover both relative positions and global UAV poses, thereby ensuring optimality at both local and global scales. In the proposed methodology, a two-stage optimization strategy is proposed to reduce optimization burden. Initially, relative Sim3 transformations among UAVs are swiftly estimated, with UWB measurements facilitating metric scale recovery in the absence of GNSS. Subsequently, a global pose optimization is performed to effectively mitigate cumulative drift. By integrating UWB, VO, and co-visibility data within this framework, both local geometric consistency and global pose accuracy are robustly maintained. Through comprehensive simulation and empirical real-world testing, we demonstrate that our system not only improves UAV positioning accuracy in challenging scenarios but also facilitates the high-quality, online integration of dense point clouds in large-scale areas. This research offers valuable contributions and practical techniques for precise, real-time map reconstruction using an autonomous UAV fleet, particularly in GNSS-denied environments.
Lin Chen 0042, Xuan Jia, Shuhui Bu, Guangming Wang 0001, Zhenyu Xia, Pengcheng Han, Xuefeng Cao
IROS3
2025 G²-Mapping: General Gaussian Mapping for Monocular, RGB-D, and LiDAR-Inertial-Visual Systems
abstract
In this paper, we introduce G2-Mapping, a novel method to comprehensively support online monocular, RGB-D, and LiDAR-Inertial-Visual systems, employing 3D gaussian points as scene representation. There are several issues when applying 3d gaussian splatting (3DGS) techniques to simultaneous localization and mapping (SLAM) 1) for monocular, the lack of depth information makes scene initialization difficult and large baseline positioning challenging; 2) differentiable rendering with respect to depth and pose has not been implemented in 3DGS, making it difficult to directly apply to the SLAM system; 3) strategy for updating the scene with incoming online frames is not present, which may lead to memory overflow. In order to overcome problems mentioned above, we formulate a mathematical derivation and propose a differentiable rendering approach that leverages both depth and color to optimize the scene and pose. We introduce a simplified odometry that provides a metric depth estimation for monocular and enhance the low-overlap scene availability. A scale consistency and uncertainty weighted optimization is further proposed to eliminates the impact of inaccurate depth prediction. Our proposed scene updating strategy effectively prevents rapid memory growth. Tracking and mapping are performed alternatively to achieve precise localization and synchronous high-fidelity map reconstruction. Extensive experiments demonstrate that our G2-Mapping surpasses feature-based SLAM in localization precision and exceeds state-of-the-art neural SLAM methods in the fidelity of view synthesis. Note to Practitioners—This paper is dedicated to tackling the efficiency challenges in multi-source SLAM and map reconstruction, aiming to generate pose and high-fidelity maps synchronously. We introduce G2-Mapping, a novel framework that leverages the power of 3D Gaussian points for scene representation, offering a universal solution for monocular, RGB-D, and LiDAR-Inertial-Visual systems. By developing a comprehensive differentiable renderer and presenting a strategy for dynamic scene updating, G2-Mapping significantly advances the state-of-the-art in localization precision and view synthesis fidelity. Although the approach is highly promising, it currently relies on the accuracy of depth prediction networks and requires further optimization for handling sparse LiDAR data. Future research will focus on enhancing these aspects, aiming to seamlessly integrate G2-Mapping into practical applications within robotics, autonomous vehicles, and augmented reality, where robust and efficient SLAM solutions are paramount.
Lin Chen 0042, Boni Hu, Jvboxi Wang, Shuhui Bu, Guangming Wang 0001, Pengcheng Han
IEEE Trans Autom. Sci. Eng.4
2024 AutoFusion: Autonomous Visual Geolocation and Online Dense Reconstruction for UAV Cluster
abstract
Real-time dense reconstruction using Unmanned Aerial Vehicle (UAV) is becoming increasingly popular in large-scale rescue and environmental monitoring tasks. However, due to the energy constraints of a single UAV, the efficiency can be greatly improved through the collaboration of multi-UAVs. Nevertheless, when faced with unknown environments or the loss of Global Navigation Satellite System (GNSS) signal, most multi-UAV SLAM systems can’t work, making it hard to construct a global consistent map. In this paper, we propose a real-time dense reconstruction system called AutoFusion for multiple UAVs, which robustly supports scenarios with lost global positioning and weak co-visibility. A method for Visual Geolocation and Matching Network (VGMN) is suggested by constructing a graph convolutional neural network as a feature extractor. It can acquire geographical location information solely through images. We also present a real-time dense reconstruction framework for multi-UAV with autonomous visual geolocation. UAV agents send images and relative positions to the ground server, which processes the data using VGMN for multi-agent geolocation optimization, including initialization, pose graph optimization, and map fusion. Extensive experiments demonstrate that our system can efficiently and stably construct large-scale dense maps in real-time with high accuracy and robustness.
Yizhu Zhang, Shuhui Bu, Yifei Dong 0008, Yu Zhang 0197, Lin Chen 0042
ICRA2
2024 High-Sample-Efficient Multiagent Reinforcement Learning for Navigation and Collision Avoidance of UAV Swarms in Multitask Environments
abstract
Multiagent reinforcement learning (MARL) algorithms have shown promise in the Internet of Things devices, such as unmanned aerial vehicle (UAV) swarms. However, the dynamic nature of large-scale swarm systems, with constantly changing numbers of agents and observed neighbors, poses challenges for MARL adaptation. Existing approaches struggle to extract meaningful features and require a substantial number of experience samples, resulting in low-sample efficiency and high-risk ratios. Moreover, these methods are effective in task-specific scenarios and fail to perform well in multitask settings. To overcome these challenges, this study proposes a high-sample efficient and scalable MARL approach for UAV swarms. The proposed approach incorporates a hypernetwork-based embedding attention (HEA) mechanism for the state representation of the policy network and a multiencoder gated transformer with a multilayer attention (MEGTrMA) mechanism for the value function. The HEA automatically generates weights for each agent to adapt to dynamic scenarios, enhancing representation ability and adaptability while reducing the cost of trial and error for improved learning efficiency. The MEGTrMA captures the contribution of each agent to the global observation, establishing long-term dependencies among them and facilitating stable policy learning in multitask scenarios. Simulation results demonstrate that the proposed method is scalable, generalizable, and high-sample efficient. Compared to learning from scratch, our method significantly reduces training time to less than one-fifth of the initial time by progressively increasing the number of UAVs and their corresponding neighbors. Additionally, the average number of collisions is reduced by an order of magnitude for large-scale UAV swarms.
Jiaming Cheng 0006, Ban Wang, Shuhui Bu
IEEE Internet Things J.4
2024 CurriculumLoc: Enhancing Cross-Domain Geolocalization Through Multistage Refinement
abstract
Visual geolocalization is a cost-effective and scalable task that involves matching one or more query images, taken at some unknown location, to a set of geotagged reference images. Existing methods, devoted to semantic features representation, evolving towards robustness to a wide variety between query and reference, including illumination and viewpoint changes, as well as scale and seasonal variations. However, practical visual geolocalization approaches need to be robust in appearance changing and extreme viewpoint variation conditions, while providing accurate global location estimates. Therefore, inspired by curriculum design, human learn general knowledge first and then delve into professional expertise. We first recognize semantic scene and then measure geometric structure. Our approach, termedCurriculumLoc, involves a delicate design of multi-stage refinement pipeline and a novel keypoint detection and description with global semantic awareness and local geometric verification. We rerank candidates and solve a particular cross-domain perspective-n-point (PnP) problem based on these keypoints and corresponding descriptors, position refinement occurs incrementally. The extensive experimental results on our collected dataset,TerraTrackand a benchmark dataset,ALTO, demonstrate that our approach results in the aforementioned desirable characteristics of a practical visual geolocalization solution. Additionally, we achieve new high recall@1 scores of 62.6% and 94.5% on ALTO, with two different distances metrics, respectively. Dataset, code and trained models are publicly available on https://github.com/npupilab/CurriculumLoc.
Boni Hu, Lin Chen 0042, Runjian Chen, Shuhui Bu, Pengcheng Han
IEEE Trans. Geosci. Remote. Sens.4
2024 MDINet: Multidomain Incremental Network for Change Detection
abstract
Traditional change detectors are ill-equipped for incremental learning (IL). Existing IL methods address the problem of catastrophic forgetting by artificially adding categories and utilizing old labels for learning supervision. Current strategies for change detection (CD) are inadequate as they fail to address a crucial aspect of the task: the constant label space throughout each training step, causing label conflicts between background-class pixels (representing unchanged regions) and changed pixels, which can lead to knowledge confusion. In this work, we revisit classical IL methods and propose an effective framework that explicitly addresses this conflict. Furthermore, we design a hierarchical distillation to ensure adequate retention of the learned features. The proposed architecture and distillation method balance acquiring new knowledge and preserving old knowledge effectively. To address the absence of datasets for IL in CD, we design a multidomain CD dataset that encompasses three distinct environments. Our proposed method demonstrates a significant improvement in the performance of IL, as measured by$\Delta _{\text {IoU}}$and$\Delta _{\text {F1}}$. Extensive experiments on this dataset show that our performance is promising compared to state-of-the-art CD methods and incremental methods.
Lean Weng, Wenqing Yang, Boni Hu, Pengcheng Han, Shaocheng Xue, Yu Zhang 0197, Shuhui Bu
IEEE Trans. Geosci. Remote. Sens.9
2023 Parallel crosschecking neural network based fault-tolerant flight parameter estimation and faulty sensor identification
Wanyong Zou, Ban Wang, Kaibo Wang, Shuhui Bu, He Shen 0001
Eng. Appl. Artif. Intell.5
2023 GCG-Net: Graph Classification Geolocation Network
abstract
Large-scale visual geolocation is a meaningful task that involves locating a query image by comparing it with images in a database and predicting the most similar image. However, the widely used training framework based on contrastive learning cannot fully utilize all data and is difficult to adapt to larger scales. At the same time, the traditional convolutional neural networks (CNNs) and vector of locally aggregated descriptors (VLADs) using aggregated features cannot fully reflect the relationship between the local features of the image. Therefore, a graph neural network (GNN) is designed as the feature extraction network, and then a training framework based on image classification is constructed. Specifically, a data grouping strategy and special loss function are designed for better training results. After training, we adopt an image retrieval strategy based on kNN for position. In addition, considering that existing datasets cannot be adapted to our requirements, two datasets are constructed for experiments, that one contains large-scale satellite images and the other fuses satellite and unmanned aerial vehicle (UAV) images. Results demonstrate that our method outperforms other common methods in both the datasets. The results demonstrate the effectiveness of our approach for UAV visual geolocation and provide ideas for future research in this field.
Yu Zhang 0197, Shuhui Bu, Boni Hu, Pengcheng Han, Lean Weng, Shaocheng Xue
IEEE Trans. Geosci. Remote. Sens.2
2022 An improved point feature-based sparse stereo vision
abstract
Abstract Since the limitation on the onboard equipment, the sparse stereo vision is becoming a suitable choice for the deployment of micro air vehicles (MAV) and small robots. However, for the point feature‐based sparse stereo, most of the current stereo algorithms ignore the similarity between feature points, so it is hard to achieve high accuracy. In addition, the problem of clustered feature distribution will still affect the performance of point feature‐based algorithms in the application. To make up for these deficiencies, the authors propose an improved features from accelerated segment test (FAST) feature detector to suppress the point detection in complex texture regions. Most importantly, the authors present a novel census transform (CT)‐based algorithm that contains two encoders ‘texture orientation’ and ‘texture gradient’ to get a more efficient census bit string for the feature point. Instead of randomly selecting pixels to calculate the bit string, we combine the texture characteristics of the census windows where feature points are located. Compared with the original CT, the processing speed of our method is improved, and the average error of our method is reduced by 18.05%. The evaluation results show the presented improved point feature‐based sparse stereo algorithm has a great value in engineering applications.
Changhao Chen, Bifeng Song, Shuhui Bu
IET Image Process.3
2022 RTSfM: Real-Time Structure From Motion for Mosaicing and DSM Mapping of Sequential Aerial Images With Low Overlap
abstract
Inspired by simultaneous localization and mapping (SLAM) style workflow, this article presented an online sequential structure from motion (SfM) solution for high-frequency video and large baseline high-resolution aerial images with high efficiency and novel precision. First, as traditional SLAM systems are not good in processing low overlap images, based on our novel hierarchical feature matching paradigm with multihomography and BoW, we proposed a robust tracking method where the relative pose and its scale are estimated separately followed by a joint optimization by considering both perspective-n-point (PnP) and epipolar constraints. Second, to further optimize the camera poses for the sparse map and dense pointcloud reconstruction, we provided a graph-based optimization with reprojection and GPS constraints, which make the camera trajectory and map georeferenced. We also incrementally generated the dense point cloud in real time from keyframes after local mapping optimization. Finally, we use a publicly available aerial image dataset with sequences of different environments, to evaluate the effectiveness of the proposed method, meanwhile, the robust performance of our solution is demonstrated with applications of high-quality aerial images mosaic and digital surface model (DSM) reconstruction in real time. Compared with the state-of-the-art SLAM and traditional SfM methods, the presented system can output large-scale high-quality ortho-mosaic and DSM in real time with the low computational cost.
Lin Chen 0042, Xishan Zhang, Shibiao Xu, Shuhui Bu, Hongkai Jiang, Pengcheng Han, Ke Li 0005
IEEE Trans. Geosci. Remote. Sens.5
2021 HMMN: Online metric learning for human re-identification via hard sample mining memory network
Pengcheng Han, Qing Li 0018, Cunbao Ma, Shibiao Xu, Shuhui Bu, Ke Li 0005
Eng. Appl. Artif. Intell.5
2021 Counting trees with point-wise supervised segmentation network
Pinmo Tong, Pengcheng Han, Suicheng Li, Shuhui Bu, Qing Li 0018, Ke Li 0005
Eng. Appl. Artif. Intell.5
2021 PL-VSCN: Patch-level vision similarity compares network for image matching
abstract
Abstract Image matching plays an important role in various computer vision tasks, such as image retrieval and loop closure detection in Simultaneous Localization and Mapping. The authors propose a discriminative patch‐based image matching method that converts the problem of whole image matching to that of local patch matching. To construct the patch representation, the Patch‐Level Vision Similarity Compare Network (PL‐VSCN) is proposed to produce the patch feature. In the image matching process, local patches that potentially contain objects within images are initially detected, and the discriminative feature of each patch is extracted based on the pre‐trained PL‐VSCN. Then, the similarities between the patch pairs are calculated to construct the similarity matrix, and the corresponding patch pairs are detected based on the mutual matching mechanism on the similarity matrix. Experimental results indicate that the proposed PL‐VSCN can generate the discriminative patch feature, which can accurately match the patch pairs with the corresponding content and distinguish those with non‐corresponding content. In addition, the comparison experiments demonstrate that the proposed image matching method outperforms existing approaches on most datasets and effectively completes the image matching task.
Xiong You, Qin Li 0005, Ke Li 0005, Anzhu Yu, Shuhui Bu
IET Comput. Vis.5
2021 Fast Georeferenced Aerial Image Stitching With Absolute Rotation Averaging and Planar- Restricted Pose Graph
abstract
Accurate digital orthophoto map generation from high-resolution aerial images is important in various applications. Compared with the existing commercial software and the current state-of-the-art mosaicing systems, a novel fast georeferenced orthophoto mosaicing framework is proposed in this study. The framework can adapt to the challenging requirements of high-accuracy orthoimage generations with relatively fast speed, even if the overlap rate is low. We provide appearance and spatial correlation-constrained fast low-overlap neighbor candidate query and matching. On the basis of GPS information, we introduce an absolute position and rotation-averaging strategy for global pose initialization, which is essential for the high convergence and efficiency of nonconvex pose optimization of every image. We also propose a planar-restricted global pose graph optimization method. The optimization is extremely efficient and robust considering that point clouds are parameterized to planes. Finally, we apply a matching graph-based exposure compensation and region reduction algorithm for large-scale and high-resolution image fusion with high efficiency and novel precision. Experimental results demonstrate that our method can achieve the state-of-the-art performance while maintaining high precision and robustness.
Guochen Liu, Shibiao Xu, Shuhui Bu, Hongkai Jiang
IEEE Trans. Geosci. Remote. Sens.4
2020 Point in: Counting Trees with Weakly Supervised Segmentation Network
abstract
For tree counting tasks, since traditional image processing methods require expensive feature engineering and are not end-to-end frameworks, this will cause additional noise and cannot be optimized overall, so this method has not been widely used in recent trends of tree counting application. Recently, many deep learning based approaches are designed for this task because of the powerful feature extracting ability. The representative way is bounding box based supervised method, but time-consuming annotations are indispensable for them. Moreover, these methods are difficult to overcome the occlusion or overlap. To solve this problem, we propose a weakly tree counting network (WTCNet) based on deep segmentation network with only point supervision. It can simultaneously complete tree counting with localization and output mask of each tree at the same time. We first adopt a novel feature extractor network (FENet) to get features of input images, and then an effective strategy is introduced to deal with different mask predictions. In the end, we propose a basic localization guidance accompany with rectification guidance to train the network. We create two different datasets and select an existing challenging plant dataset to evaluate our method on three different tasks. Experimental results show the good performance improvement of our method compared with other existing methods. Further study shows that our method has great potential to reduce human labor and provide effective ground-truth masks and the results show the superiority of our method over the advanced methods.
Pinmo Tong, Xishan Zhang, Pengcheng Han, Shuhui Bu
ICPR4
2020 DenseFusion: Large-Scale Online Dense Pointcloud and DSM Mapping for UAVs
abstract
With the rapidly developing unmanned aerial vehicles, the requirements of generating maps efficiently and quickly are increasing. To realize online mapping, we develop a real-time dense mapping framework named DenseFusion which can incrementally generates dense geo-referenced 3D point cloud, digital orthophoto map (DOM) and digital surface model (DSM) from sequential aerial images with optional GPS information. The proposed method works in real-time on standard CPUs even for processing high resolution images. Based on the advanced monocular SLAM, our system first estimates appropriate camera poses and extracts effective keyframes, and next constructs virtual stereo-pair from consecutive frame to generate pruned dense 3D point clouds; then a novel realtime DSM fusion method is proposed which can incrementally process dense point cloud. Finally, a high efficiency visualization system is developed to adopt dynamic levels of detail (LoD) method, which makes it render dense point cloud and DSM smoothly. The performance of the proposed method is evaluated through qualitative and quantitative experiments. The results indicate that compared to traditional structure from motion based approaches, the presented framework is able to output both large-scale high-quality DOM and DSM in real-time with low computational cost.
Lin Chen 0042, Shibiao Xu, Shuhui Bu, Pengcheng Han
IROS4
2020 Change detection in images using shape-aware siamese convolutional network
Suicheng Li, Pengcheng Han, Shuhui Bu, Pinmo Tong, Qing Li 0018, Ke Li 0005
Eng. Appl. Artif. Intell.3
2020 Mask-CDNet: A mask based pixel change detection network
Shuhui Bu, Qing Li 0018, Pengcheng Han, Pengyu Leng, Ke Li 0005
Neurocomputing1
2019 GSLAM: A General SLAM Framework and Benchmark
abstract
SLAM technology has recently seen many successes and attracted the attention of high-technological companies. However, how to unify the interface of existing or emerging algorithms, and effectively perform benchmark about the speed, robustness and portability are still problems. In this paper, we propose a novel SLAM platform named GSLAM, which not only provides evaluation functionality, but also supplies useful toolkit for researchers to quickly develop their SLAM systems. Our core contribution is an universal, cross-platform and full open-source SLAM interface for both research and commercial usage, which is aimed to handle interactions with input dataset, SLAM implementation, visualization and applications in an unified framework. Through this platform, users can implement their own functions for better performance with plugin form and further boost the application to practical usage of the SLAM.
Shibiao Xu, Shuhui Bu, Hongkai Jiang, Pengcheng Han
ICCV3
2019 TerrainFusion: Real-time Digital Surface Model Reconstruction based on Monocular SLAM
abstract
This paper presents an algorithm which can generate live digtial surface model (DSM) during the flight based on simultaneous localization and mapping (SLAM). We process the keyframe which is output by a monocular SLAM system to generate a local DSM, and fuse the local DSM to the global tiled DSM incrementally. During the local DSM generation, a local digital elevation model (DEM) is estimated by projecting the filtered 2D Delaunay mesh to a 3D mesh, and a local orthomosaic is obtained by projecting triangle image patches onto a 2D mesh. During the DSM fusion, both the local DEM and orthomosaic are split into tiles and fused to the global tiled DEM and orthomosaic respectively with multiband algorithm. Both the efficient DSM generation and fusion algorithms contribute to achieving a real-time reconstruction. Qualitative and quantitative experiments on a public aerial image dataset with different scenarios are performed to validate the effectiveness of the proposed method. Compared with traditional structure from motion (SfM) based approaches, the presented system is able to output both large-scale high-quality DEM and orthomosaic in real-time with low computational cost.
Pengcheng Han, Shuhui Bu
IROS5
2019 Aerial image change detection using dual regions of interest networks
Pengcheng Han, Cunbao Ma, Qing Li 0018, Pengyu Leng, Shuhui Bu, Ke Li 0005
Neurocomputing5
2019 Unsupervised Learning of 3-D Local Features From Raw Voxels Based on a Novel Permutation Voxelization Strategy
abstract
Effective 3-D local features are significant elements for 3-D shape analysis. Existing hand-crafted 3-D local descriptors are effective but usually involve intensive human intervention and prior knowledge, which burdens the subsequent processing procedures. An alternative resorts to the unsupervised learning of features from raw 3-D representations via popular deep learning models. However, this alternative suffers from several significant unresolved issues, such as irregular vertex topology, arbitrary mesh resolution, orientation ambiguity on the 3-D surface, and rigid and slightly nonrigid transformation invariance. To tackle these issues, we propose an unsupervised 3-D local feature learning framework based on a novel permutation voxelization strategy to learn high-level and hierarchical 3-D local features from raw 3-D voxels. Specifically, the proposed strategy first applies a novel voxelization which discretizes each 3-D local region with irregular vertex topology and arbitrary mesh resolution into regular voxels, and then, a novel permutation is applied to permute the voxels to simultaneously eliminate the effect of rotation transformation and orientation ambiguity on the surface. Based on the proposed strategy, the permuted voxels can fully encode the geometry and structure of each local region in regular, sparse, and binary vectors. These voxel vectors are highly suitable for the learning of hierarchical common surface patterns by stacked sparse autoencoder with hierarchical abstraction and sparse constraint. Experiments are conducted on three aspects for evaluating the learned local features: 1) global shape retrieval; 2) partial shape retrieval; and 3) shape correspondence. The experimental results show that the learned local features outperform the other state-of-the-art 3-D shape descriptors.
Zhizhong Han, Zhenbao Liu, Junwei Han 0001, Chi-Man Vong, Shuhui Bu, C. L. Philip Chen
IEEE Trans. Cybern.5
2018 Multi-modal feature fusion for geographic image annotation
Ke Li 0005, Changqing Zou, Shuhui Bu, Yun Liang 0003, Jian Zhang 0026, Minglun Gong
Pattern Recognit.3
2018 Rotation-Insensitive and Context-Augmented Object Detection in Remote Sensing Images
abstract
Most of the existing deep-learning-based methods are difficult to effectively deal with the challenges faced for geospatial object detection such as rotation variations and appearance ambiguity. To address these problems, this paper proposes a novel deep-learning-based object detection framework including region proposal network (RPN) and local-contextual feature fusion network designed for remote sensing images. Specifically, the RPN includes additional multiangle anchors besides the conventional multiscale and multiaspect-ratio ones, and thus can deal with the multiangle and multiscale characteristics of geospatial objects. To address the appearance ambiguity problem, we propose a double-channel feature fusion network that can learn local and contextual properties along two independent pathways. The two kinds of features are later combined in the final layers of processing in order to form a powerful joint representation. Comprehensive evaluations on a publicly available ten-class object detection data set demonstrate the effectiveness of the proposed method.
Ke Li 0005, Gong Cheng 0003, Shuhui Bu, Xiong You
IEEE Trans. Geosci. Remote. Sens.3
2018 Automatic Building Rooftop Extraction From Aerial Images via Hierarchical RGB-D Priors
abstract
Accurate building rooftop extraction from high-resolution aerial images is of crucial importance in a wide range of applications. Owing to the varying appearance and large-scale range of scene objects, especially for building rooftops in different scales and heights, single-scale or individual prior-based extraction technique is insufficient in pursuing efficient, generic, and accurate extraction results. The trend toward integrating multiscale or several cue techniques appears to be the best way; thus, such integration is the focus of this paper. We first propose a novel salient rooftop detector integrating four correlative RGB-D priors (depth cue, uniqueness prior, shape prior, and transition surface prior) for improved rooftop extraction to address the preceding complex issues mentioned. Then, these correlative cues are computed from image layers created by our multilevel segmentation and further fused into the state-of-the-art high-order conditional random field (CRF) framework to locate the rooftop. Finally, an iterative optimization strategy is applied for high-quality solving, which can robustly handle varying appearance of building rooftops. Performance evaluations in the SZTAKI-INRIA benchmark data sets show that our method outperforms the traditional color-based algorithm and the original high-order CRF algorithm and its variants. The proposed algorithm is also evaluated and found to produce consistently satisfactory results for various large-scale, real-world data sets.
Shibiao Xu, Xingjia Pan, Er Li, Baoyuan Wu, Shuhui Bu, Weiming Dong, Shiming Xiang, Xiaopeng Zhang 0001
IEEE Trans. Geosci. Remote. Sens.5
2018 Deep Spatiality: Unsupervised Learning of Spatially-Enhanced Global and Local 3D Features by Deep Neural Network With Coupled Softmax
abstract
The discriminability of Bag-of-Words representations can be increased via encoding the spatial relationship among virtual words on 3D shapes. However, this encoding task involves several issues, including arbitrary mesh resolutions, irregular vertex topology, orientation ambiguity on 3D surface, invariance to rigid and non-rigid shape transformations. To address these issues, a novel unsupervised spatial learning framework based on deep neural network, deep spatiality (DS), is proposed. Specifically, DS employs two novel components: spatial context extractor and deep context learner. Spatial context extractor extracts the spatial relationship among virtual words in a local region into a raw spatial representation. Along a consistent circular direction, a directed circular graph is constructed to encode relative positions between pairwise virtual words in each face ring into a relative spatial matrix. By decomposing each relative spatial matrix using SVD, the raw spatial representation is formed, from which deep context learner conducts unsupervised learning of global and local features. Deep context learner is a deep neural network with a novel model structure to adapt the proposed coupled softmax layer, which encodes not only the discriminative information among local regions but also the one among global shapes. Experimental results show that DS outperforms state-of-the-art methods.
Zhizhong Han, Zhenbao Liu, Chi-Man Vong, Yu-Shen Liu, Shuhui Bu, Junwei Han 0001, C. L. Philip Chen
IEEE Trans. Image Process.5
2017 3D shape recognition and retrieval based on multi-modality deep learning
Shuhui Bu, Pengcheng Han, Zhenbao Liu, Ke Li 0005
Neurocomputing1
2017 Semi-direct tracking and mapping with RGB-D camera for MAV
Shuhui Bu, Ke Li 0005, Gong Cheng 0003, Zhenbao Liu
Multim. Tools Appl.1
2017 Human Motion Tracking by Multiple RGBD Cameras
abstract
The advent of low-cost depth cameras, such as the Microsoft Kinect in the consumer market, has made many indoor applications and games based on motion tracking available to the everyday user. However, it is a large challenge to track human motion via such a camera because of its low-quality images, missing depth values, and noise. In this paper, we propose a novel human motion capture method based on a cooperative structure of multiple low-cost RGBD cameras, which can effectively avoid these problems. This structure can also manage the problem of body occlusions that appears when a single camera is used. Moreover, the whole process does not require training data, which makes this approach easily deployed and reduces operation time. We use the color image, depth image, and point cloud acquired in each view as the data source, and an initial pose is extracted in our optimization framework by aligning multiple point clouds from different cameras. The pose is dynamically updated by combining a filtering approach with a Markov model to estimate new poses in video streams. To verify the efficiency and robustness of our approach, we capture a wide variety of human actions via three cameras in indoor scenes and compare the tracking results of the proposed method to those of the current state-of-the-art methods. Moreover, our system is tested on more complex situations, in which multiple humans move within a scene, possibly occluding each other to some extent. The actions of multiple humans are tracked simultaneously, which would assist group behavior analysis.
Zhenbao Liu, Jinxin Huang, Junwei Han 0001, Shuhui Bu, Jianfeng Lv
IEEE Trans. Circuits Syst. Video Technol.4
2017 Template Deformation-Based 3-D Reconstruction of Full Human Body Scans From Low-Cost Depth Cameras
abstract
Full human body shape scans provide valuable data for a variety of applications including anthropometric surveying, clothing design, human-factors engineering, health, and entertainment. However, the high price, large volume, and difficulty of operating professional 3-D scanners preclude their use in home entertainment. Recently, portable low-cost red green blue-depth cameras such as the Kinect have become popular for computer vision tasks. However, the infrared mechanism of this type of camera leads to noisy and incomplete depth images. We construct a stereo full-body scanning environment composed of multiple depth cameras and propose a novel registration algorithm. Our algorithm determines a segment constrained correspondence for two neighboring views, integrating them using rigid transformation. Furthermore, it aligns all of the views based on uniform error distribution. The generated 3-D mesh model is typically sparse, noisy, and even with holes, which makes it lose surface details. To address this, we introduce a geometric and topological fitting prior in the form of a professionally designed high-resolution template model. We formulate a template deformation optimization problem to fit the high-resolution model to the low-quality scan. Its solution overcomes the obstacles posed by different poses, varying body details, and surface noise. The entire process is free of body and template markers, fully automatic, and achieves satisfactory reconstruction results.
Zhenbao Liu, Jinxin Huang, Shuhui Bu, Junwei Han 0001, Xuelong Li 0001
IEEE Trans. Cybern.3
2017 Capturing High-Discriminative Fault Features for Electronics-Rich Analog System via Deep Learning
abstract
Fault detection and isolation (FDI) is very difficult for electronics-rich analog systems due to its sophisticated mechanism and variable operational conditions. Traditionally, FDI in such systems is done through the monitoring of deviation of output signals in voltage or current at system level, which commonly arises from the degradation of one or more critical components. Therefore, FDI can be transformed to a multiclass classification task given the extracted features of the output signals in voltage or current of the circuit. Traditional feature extraction on the circuit output is mostly based on time-domain, frequency-domain, or time-frequency signal processing, which collapse high-dimensional raw signals into a lower dimensional feature set. Such low-dimensional feature set usually suffers from information loss so as to affect the accuracy of the later fault diagnosis. In order to retain as much information as possible, deep learning is proposed which employs a hierarchical structure to capture the different levels of semantic representations of the signals. In this paper, a novel fault diagnostic application of Gaussian-Bernoulli deep belief network (GB-DBN) for electronics-rich analog systems is developed which can more effectively capture the high-order semantic features within the raw output signals. The novel fault diagnosis is validated experimentally on two typical analog filter circuits. Experimental results show the fault diagnosis based on GB-DBN is with superior diagnostic performance than the traditional feature extraction methods.
Zhenbao Liu, Chi-Man Vong, Shuhui Bu, Junwei Han 0001
IEEE Trans. Ind. Informatics4
2017 BoSCC: Bag of Spatial Context Correlations for Spatially Enhanced 3D Shape Representation
abstract
Highly discriminative 3D shape representations can be formed by encoding the spatial relationship among virtual words into the Bag of Words (BoW) method. To achieve this challenging task, several unresolved issues in the encoding procedure must be overcome for 3D shapes, including: 1) arbitrary mesh resolution; 2) irregular vertex topology; 3) orientation ambiguity on the 3D surface; and 4) invariance to rigid and non-rigid shape transformations. In this paper, a novel spatially enhanced 3D shape representation called bag of spatial context correlations (BoSCCs) is proposed to address all these issues. Adopting a novel local perspective, BoSCC is able to describe a 3D shape by an occurrence frequency histogram of spatial context correlation patterns, which makes BoSCC become more compact and discriminative than previous global perspective-based methods. Specifically, the spatial context correlation is proposed to simultaneously encode the geometric and spatial information of a 3D local region by the correlation among spatial contexts of vertices in that region, which effectively resolves the aforementioned issues. The spatial context of each vertex is modeled by Markov chains in a multi-scale manner, which thoroughly captures the spatial relationship by the transition probabilities of intra-virtual words and the ones of inter-virtual words. The high discriminability and compactness of BoSCC are effective for classification and retrieval, especially in the scenarios of limited samples and partial shape retrieval. Experimental results show that BoSCC outperforms the state-of-the-art spatially enhanced BoW methods in three common applications: global shape retrieval, shape classification, and partial shape retrieval.
Zhizhong Han, Zhenbao Liu, Chi-Man Vong, Yu-Shen Liu, Shuhui Bu, Junwei Han 0001, C. L. Philip Chen
IEEE Trans. Image Process.5
2017 Mesh Convolutional Restricted Boltzmann Machines for Unsupervised Learning of Features With Structure Preservation on 3-D Meshes
abstract
Discriminative features of 3-D meshes are significant to many 3-D shape analysis tasks. However, handcrafted descriptors and traditional unsupervised 3-D feature learning methods suffer from several significant weaknesses: 1) the extensive human intervention is involved; 2) the local and global structure information of 3-D meshes cannot be preserved, which is in fact an important source of discriminability; 3) the irregular vertex topology and arbitrary resolution of 3-D meshes do not allow the direct application of the popular deep learning models; 4) the orientation is ambiguous on the mesh surface; and 5) the effect of rigid and nonrigid transformations on 3-D meshes cannot be eliminated. As a remedy, we propose a deep learning model with a novel irregular model structure, called mesh convolutional restricted Boltzmann machines (MCRBMs). MCRBM aims to simultaneously learn structure-preserving local and global features from a novel raw representation, local function energy distribution. In addition, multiple MCRBMs can be stacked into a deeper model, called mesh convolutional deep belief networks (MCDBNs). MCDBN employs a novel local structure preserving convolution (LSPC) strategy to convolve the geometry and the local structure learned by the lower MCRBM to the upper MCRBM. LSPC facilitates resolving the challenging issue of the orientation ambiguity on the mesh surface in MCDBN. Experiments using the proposed MCRBM and MCDBN were conducted on three common aspects: global shape retrieval, partial shape retrieval, and shape correspondence. Results show that the features learned by the proposed methods outperform the other state-of-the-art 3-D shape features.
Zhizhong Han, Zhenbao Liu, Junwei Han 0001, Chi-Man Vong, Shuhui Bu, C. L. Philip Chen
IEEE Trans. Neural Networks Learn. Syst.5
2016 Map2DFusion: Real-time incremental UAV image mosaicing based on monocular SLAM
abstract
In this paper we present a real-time approach to stitch large-scale aerial images incrementally. A monocular SLAM system is used to estimate camera position and attitude, and meanwhile 3D point cloud map is generated. When GPS information is available, the estimated trajectory is transformed to WGS84 coordinates after time synchronized automatically. Therefore, the output orthoimage retains global coordinates without ground control points. The final image is fused and visualized instantaneously with a proposed adaptive weighted multiband algorithm. To evaluate the effectiveness of the proposed method, we create a publicly available aerial image dataset with sequences of different environments. The experimental results demonstrate that our system is able to achieve high efficiency and quality compared to state-of-the-art methods. In addition, we share the code on the website with detailed introduction and results.
Shuhui Bu, Zhenbao Liu
IROS1
2016 Place recognition based on deep feature and adaptive weighting of similarity matrix
Qin Li 0005, Ke Li 0005, Xiong You, Shuhui Bu, Zhenbao Liu
Neurocomputing4
2016 Scene parsing using inference Embedded Deep Networks
Shuhui Bu, Pengcheng Han, Zhenbao Liu, Junwei Han 0001
Pattern Recognit.1
2016 Unsupervised 3D Local Feature Learning by Circle Convolutional Restricted Boltzmann Machine
abstract
Extracting local features from 3D shapes is an important and challenging task that usually requires carefully designed 3D shape descriptors. However, these descriptors are hand-crafted and require intensive human intervention with prior knowledge. To tackle this issue, we propose a novel deep learning model, namely circle convolutional restricted Boltzmann machine (CCRBM), for unsupervised 3D local feature learning. CCRBM is specially designed to learn from raw 3D representations. It effectively overcomes obstacles such as irregular vertex topology, orientation ambiguity on the 3D surface, and rigid or slightly non-rigid transformation invariance in the hierarchical learning of 3D data that cannot be resolved by the existing deep learning models. Specifically, by introducing the novel circle convolution, CCRBM holds a novel ring-like multi-layer structure to learn 3D local features in a structure preserving manner. Circle convolution convolves across 3D local regions via rotating a novel circular sector convolution window in a consistent circular direction. In the process of circle convolution, extra points are sampled in each 3D local region and projected onto the tangent plane of the center of the region. In this way, the projection distances in each sector window are employed to constitute a novel local raw 3D representation called projection distance distribution (PDD). In addition, to eliminate the initial location ambiguity of a sector window, the Fourier transform modulus is used to transform the PDD into the Fourier domain, which is then conveyed to CCRBM. Experiments using the learned local features are conducted on three aspects: global shape retrieval, partial shape retrieval, and shape correspondence. The experimental results show that the learned local features outperform other state-of-the-art 3D shape descriptors.
Zhizhong Han, Zhenbao Liu, Junwei Han 0001, Chi-Man Vong, Shuhui Bu, Xuelong Li 0001
IEEE Trans. Image Process.5
2016 Human-Centered Saliency Detection
abstract
We introduce a new concept for detecting the saliency of 3-D shapes, that is, human-centered saliency (HCS) detection on the surface of shapes, whereby a given shape is analyzed not based on geometric or topological features directly obtained from the shape itself, but by studying how a human uses the object. Using virtual agents to simulate the ways in which humans interact with objects helps to understand shapes and detect their salient parts in relation to their functions. HCS detection is less affected by inconsistencies between the geometry or topology of the analyzed 3-D shapes. The potential benefit of the proposed method is that it is adaptable to variable shapes with the same semantics, as well as being robust against a geometrical and topological noise. Given a 3-D shape, its salient part is detected by automatically selecting a corresponding agent and making them interact with each other. Their adaption and alignment depend on an optimization framework and a training process. We demonstrate the detected salient parts for different types of objects together with the stability thereof. The salient parts can be used for important vision tasks, such as 3-D shape retrieval.
Zhenbao Liu, Xiao Wang 0025, Shuhui Bu
IEEE Trans. Neural Networks Learn. Syst.3
2015 Local deep feature learning framework for 3D shape
Shuhui Bu, Pengcheng Han, Zhenbao Liu, Junwei Han 0001
Comput. Graph.1
2015 Indirect shape analysis for 3D shape retrieval
Zhenbao Liu, Caili Xie, Shuhui Bu, Xiao Wang 0025, Junwei Han 0001, Hao (Richard) Zhang
Comput. Graph.3
2015 Locality-constrained sparse patch coding for 3D shape retrieval
Zhenbao Liu, Shuhui Bu, Junwei Han 0001
Neurocomputing2
2015 A coarse-to-fine model for airport detection from remote sensing images using target-oriented visual saliency and CRF
Xiwen Yao, Junwei Han 0001, Lei Guo 0002, Shuhui Bu, Zhenbao Liu
Neurocomputing4
2015 Weakly Supervised Learning for Target Detection in Remote Sensing Images
abstract
In this letter, we develop a novel framework of leveraging weakly supervised learning techniques to efficiently detect targets from remote sensing images, which enables us to reduce the tedious manual annotation for collecting training data while maintaining the detection accuracy to large extent. The proposed framework consists of a weakly supervised training procedure to yield the detectors and an effective scheme to detect targets from testing images. Comprehensive evaluations on three benchmarks which have different spatial resolutions and contain different types of targets as well as the comparisons with traditional supervised learning schemes demonstrate the efficiency and effectiveness of the proposed framework.
Dingwen Zhang, Junwei Han 0001, Gong Cheng 0003, Zhenbao Liu, Shuhui Bu, Lei Guo 0002
IEEE Geosci. Remote. Sens. Lett.5
2015 3D real human reconstruction via multiple low-cost depth cameras
Zhenbao Liu, Hongliang Qin, Shuhui Bu, Meng Yan 0006, Jinxin Huang, Junwei Han 0001
Signal Process.3
2015 Effective and Efficient Midlevel Visual Elements-Oriented Land-Use Classification Using VHR Remote Sensing Images
abstract
Land-use classification using remote sensing images covers a wide range of applications. With more detailed spatial and textural information provided in very high resolution (VHR) remote sensing images, a greater range of objects and spatial patterns can be observed than ever before. This offers us a new opportunity for advancing the performance of land-use classification. In this paper, we first introduce an effective midlevel visual elementsoriented land-use classification method based on “partlets,” which are a library of pretrained part detectors used for midlevel visual elements discovery. Taking advantage of midlevel visual elements rather than low-level image features, a partlets-based method represents images by computing their responses to a large number of part detectors. As the number of part detectors grows, a main obstacle to the broader application of this method is its computational cost. To address this problem, we next propose a novel framework to train coarse-to-fine shared intermediate representations, which are termed “sparselets,” from a large number of pretrained part detectors. This is achieved by building a single-hidden-layer autoencoder and a single-hidden-layer neural network with an L0-norm sparsity constraint, respectively. Comprehensive evaluations on a publicly available 21-class VHR landuse data set and comparisons with state-of-the-art approaches demonstrate the effectiveness and superiority of this paper.
Gong Cheng 0003, Junwei Han 0001, Lei Guo 0002, Zhenbao Liu, Shuhui Bu, Jinchang Ren
IEEE Trans. Geosci. Remote. Sens.5
2015 3D shape creation by style transfer
Zhizhong Han, Zhenbao Liu, Junwei Han 0001, Shuhui Bu
Vis. Comput.4
2014 High-level semantic feature for 3D shape based on deep belief networks
abstract
Deep learning has emerged as a powerful technique to extract high-level features from low-level information, which shows that hierarchical representation can be easily achieved. However, applying deep learning into 3D shape is still a challenge. In this paper, we propose a novel high-level feature learning method for 3D shape retrieval based on deep learning. In this framework, the low-level 3D shape descriptors are first encoded into visual bag-of-words, and then highlevel shape features are generated via deep belief network, which facilitates a good semantic preserving ability for the tasks of shape classification and retrieval. Experiments on 3D shape recognition and retrieval demonstrate the superior performance of the proposed method in comparison to the state-of-the-art methods.
Zhenbao Liu, Shaoguang Chen, Shuhui Bu, Ke Li 0005
ICME3
2014 A 3D Fingertips Detecting and Tracking Algorithm based on the Sliding Window
abstract
In this paper, we propose a 3D finger detecting and tracking algorithm (3dFDT) based on a sliding window for RGB-D sequences. Microsoft Kinect is utilized as a 3D depth camera, which provides RGB images and depth data simultaneously. However the depth data have large noises when the background is white. We are focusing on tracking the up or down movements of the fingertips in order to apply in some applications such as playing the piano on the table or " in the air". First, the skin statistical ellipse model is applied to detect the hand skin region, and the pseudo-contours are removed to get the interested hand region. Then, the region center is considered as the center of the palm. Next, the fingertips are located precisely using the convex defect detection based on the refined hand contours. Furthermore, the depth information is aligned to the fingertips in the RGB channels. Finally, the sliding window strategy is employed to stabilize the depth information due to large noises from the depth space especially when dealing with the white surface as backgrounds. The experimental results show that our proposed method is effective, and it can be applied in the real-time applications for non-contact interactions.
Wenjing Qiao, Cailiang Kuang, Zhenbao Liu, Shuhui Bu, Junwei Han 0001
ACM Multimedia5
2014 Sparse Patch Coding for 3D Model Retrieval
Zhenbao Liu, Shuhui Bu, Junwei Han 0001
MMM (2)2
2014 Spectral Classification of 3D Articulated Shapes
Zhenbao Liu, Shuhui Bu
MMM (2)3
2014 Automatic 3D Indoor Scene Updating with RGBD Cameras
abstract
Abstract Since indoor scenes are frequently changed in daily life, such as re‐layout of furniture, the 3D reconstructions for them should be flexible and easy to update. We present an automatic 3D scene update algorithm to indoor scenes by capturing scene variation with RGBD cameras. We assume an initial scene has been reconstructed in advance in manual or other semi‐automatic way before the change, and automatically update the reconstruction according to the newly captured RGBD images of the real scene update. It starts with an automatic segmentation process without manual interaction, which benefits from accurate labeling training from the initial 3D scene. After the segmentation, objects captured by RGBD camera are extracted to form a local updated scene. We formulate an optimization problem to compare to the initial scene to locate moved objects. The moved objects are then integrated with static objects in the initial scene to generate a new 3D scene. We demonstrate the efficiency and robustness of our approach by updating the 3D scene of several real‐world scenes.
Zhenbao Liu, Sicong Tang, Weiwei Xu 0003, Shuhui Bu, Junwei Han 0001, Kun Zhou 0001
Comput. Graph. Forum4
2014 Learning High-Level Feature by Deep Belief Networks for 3-D Model Retrieval and Recognition
abstract
3-D shape analysis has attracted extensive research efforts in recent years, where the major challenge lies in designing an effective high-level 3-D shape feature. In this paper, we propose a multi-level 3-D shape feature extraction framework by using deep learning. The low-level 3-D shape descriptors are first encoded into geometric bag-of-words, from which middle-level patterns are discovered to explore geometric relationships among words. After that, high-level shape features are learned via deep belief networks, which are more discriminative for the tasks of shape classification and retrieval. Experiments on 3-D shape recognition and retrieval demonstrate the superior performance of the proposed method in comparison to the state-of-the-art methods.
Shuhui Bu, Zhenbao Liu, Junwei Han 0001, Rongrong Ji
IEEE Trans. Multim.1
2014 Shift-invariant ring feature for 3D shape
Shuhui Bu, Pengcheng Han, Zhenbao Liu, Ke Li 0005, Junwei Han 0001
Vis. Comput.1
2013 Superpixel segmentation based structural scene recognition
abstract
This paper presents a novel structural model based scene recognition method. In order to resolve regular grid image division methods which cause low content discriminability for scene recognition in previous methods, we partition an image into a pre-defined set of regions by superpixel segmentation. And then classification is modelled by introducing a structural model which has the capability of organizing unordered features of image patches. In the implementation, CENTRIST which is robust to scene recognition is used as original image feature, and bag-of-words representation is used to capture the local appearances of an image. In addition, we incorporate adjacent superpixel's differences as edge features. Our models are trained using structural SVM. Two state-of-the-art scene datasets are adopted to evaluate the proposed method. The experiment results show that the recognition accuracy is significantly improved by the proposed method.
Shuhui Bu, Zhenbao Liu, Junwei Han 0001
ACM Multimedia1
2013 New evaluation metrics for mesh segmentation
Zhenbao Liu, Sicong Tang, Shuhui Bu, Hao (Richard) Zhang
Comput. Graph.3
2013 A Survey on Partial Retrieval of 3D Shapes
Zhenbao Liu, Shuhui Bu, Kun Zhou 0001, Shuming Gao, Junwei Han 0001
J. Comput. Sci. Technol.2
2012 Semi-supervised adaptive parzen Gentleboost algorithm for fault diagnosis
Chengliang Li, Zhongsheng Wang, Shuhui Bu, Zhenbao Liu
ICPR3
2012 Evaluating user's energy consumption using kinect based skeleton tracking
abstract
By utilizing the dataset provided by 3DLife/Huawei Challenge of ACM Multimedia, we propose a refreshing application that automatically evaluates player's energy consumption in gaming scenarios by a model with tracked skeleton, which may help users to know their exercise effects and even diet or reduce their weights. We develop a program to compute the energy consumption in real time by analyzing data captured from Microsoft Kinect, and also give a cue in the dynamic interaction. We model 3D human skeleton by joining different body parts with 15 nodes, and decompose player action into rigid body motions of these parts. Amount of energy consumed in the action is calculated as the sum of powers required to overcome gravity of each part. Experimental results show that instantaneous and total energy consumption of different dancers can be stably calculated. The hardware system is based on low-price Kinect, and easily accepted by users. The proposed application also provides a quantitative approach which help users to control their dining and exercise intensity.
Zhenbao Liu, Sicong Tang, Hongliang Qin, Shuhui Bu
ACM Multimedia4