EDBT 2026 Demo / reviewers in the wild / expert
Minghua Liu
dblp:28/8907
· DBLP profile ↗
28ranked-venue papers
9as first author
20since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 18 · 9 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 16 · 6 first-author · 11 since 2021Human-computer interaction and ubiquitous computing · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Systems, architecture and hardware · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Dual-branch spatio-temporal enhancement network-based human action recognition in UAV scenarios
Minghua Liu, Xinyi Mao |
Expert Syst. Appl. | 2 |
| 2026 | GLDS-GCN: Dual-Stream graph convolutional network for skeleton-based abnormal action recognition
Minghua Liu, Xinyi Mao |
Knowl. Based Syst. | 2 |
| 2025 | FreeArt3D: Training-Free Articulated Object Generation using 3D DiffusionabstractArticulated 3D objects are central to many applications in robotics, AR/VR, and animation. Recent approaches to modeling such objects either rely on optimization-based reconstruction pipelines that require dense-view supervision or on feed-forward generative models that produce coarse geometric approximations and often overlook surface texture. In contrast, open-world 3D generation of static objects has achieved remarkable success, especially with the advent of native 3D diffusion models such as Trellis. However, extending these methods to articulated objects by training native 3D diffusion models poses significant challenges. In this work, we present FreeArt3D, a training-free framework for articulated 3D object generation. Instead of training a new model on limited articulated data, FreeArt3D repurposes a pre-trained static 3D diffusion model (e.g., Trellis) as a powerful shape prior. It extends Score Distillation Sampling (SDS) into the 3D-to-4D domain by treating articulation as an additional generative dimension. Given a few images captured in different articulation states, FreeArt3D jointly optimizes the object’s geometry, texture, and articulation parameters—without requiring task-specific training or access to large-scale articulated datasets. Our method generates high-fidelity geometry and textures, accurately predicts underlying kinematic structures, and generalizes well across diverse object categories. Despite following a per-instance optimization paradigm, FreeArt3D completes in minutes and significantly outperforms prior state-of-the-art approaches in both quality and versatility. Code for this paper is at https://github.com/CzzzzH/FreeArt3D. Chuhao Chen 0003, Isabella Liu, Xinyue Wei, Hao Su 0001, Minghua Liu |
SIGGRAPH Asia | 5 |
| 2025 | PartUV: Part-Based UV Unwrapping of 3D MeshesabstractUV unwrapping flattens 3D surfaces to 2D with minimal distortion, often requiring the complex surface to be decomposed into multiple charts. Although extensively studied, existing UV unwrapping methods frequently struggle with AI-generated meshes, which are typically noisy, bumpy, and poorly conditioned. These methods often produce highly fragmented charts and suboptimal boundaries, introducing artifacts and hindering downstream tasks. We introduce PartUV, a part-based UV unwrapping pipeline that generates significantly fewer, part-aligned charts while maintaining low distortion. Built on top of a recent learning-based part decomposition method PartField, PartUV combines high-level semantic part decomposition with novel geometric heuristics in a top-down recursive framework. It ensures each chart’s distortion remains below a user-specified threshold while minimizing the total number of charts. The pipeline integrates and extends parameterization and packing algorithms, incorporates dedicated handling of non-manifold and degenerate meshes, and is extensively parallelized for efficiency. Evaluated across four diverse datasets—including man-made, CAD, AI-generated, and Common Shapes—PartUV outperforms existing tools and recent neural methods in chart count and seam length, achieves comparable distortion, exhibits high success rates on challenging meshes, and enables new applications like part-specific multi-tiles packing. Code for this paper is at https://github.com/EricWang12/PartUV. Zhaoning Wang, Xinyue Wei, Ruoxi Shi, Xiaoshuai Zhang, Hao Su 0001, Minghua Liu |
SIGGRAPH Asia | 6 |
| 2025 | LARM: A Large Articulated Object Reconstruction ModelabstractModeling 3D articulated objects with realistic geometry, textures, and kinematics is essential for a wide range of applications. However, existing optimization-based reconstruction methods often require dense multi-view inputs and expensive per-instance optimization, limiting their scalability. Recent feedforward approaches offer faster alternatives but frequently produce coarse geometry, lack texture reconstruction, and rely on brittle, complex multi-stage pipelines. We introduce LARM, a unified feedforward framework that reconstructs 3D articulated objects from sparse-view images by jointly recovering detailed geometry, realistic textures, and accurate joint structures. LARM extends LVSM—a recent novel view synthesis (NVS) approach for static 3D objects—into the articulated setting by jointly reasoning over camera pose and articulation variation using a transformer-based architecture, enabling scalable and accurate novel view synthesis. In addition, LARM generates auxiliary outputs such as depth maps and part masks to facilitate explicit 3D mesh extraction and joint estimation. Our pipeline eliminates the need for dense supervision and supports high-fidelity reconstruction across diverse object categories. Extensive experiments demonstrate that LARM outperforms state-of-the-art methods in both novel view and state synthesis as well as 3D articulated object reconstruction, generating high-quality meshes that closely adhere to the input images. Code for this paper is at https://github.com/sylviayuan-sy/LARM. Sylvia Yuan, Ruoxi Shi, Xinyue Wei, Xiaoshuai Zhang, Hao Su 0001, Minghua Liu |
SIGGRAPH Asia | 6 |
| 2025 | PaMO: Parallel Mesh Optimization for Intersection-Free Low-Poly Modeling on the GPUabstractAbstract Reducing the triangle count in complex 3D models is a basic geometry preprocessing step in graphics pipelines such as efficient rendering and interactive editing. However, most existing mesh simplification methods exhibit a few issues. Firstly, they often lead to self‐intersections during decimation, a major issue for applications such as 3D printing and soft‐body simulation. Second, to perform simplification on a mesh in the wild, one would first need to perform re‐meshing, which often suffers from surface shifts and losses of sharp features. Finally, existing re‐meshing and simplification methods can take minutes when processing large‐scale meshes, limiting their applications in practice. To address the challenges, we introduce a novel GPU‐based mesh optimization approach containing three key components: (1) a parallel re‐meshing algorithm to turn meshes in the wild into watertight, manifold, and intersection‐free ones, and reduce the prevalence of poorly shaped triangles; (2) a robust parallel simplification algorithm with intersection‐free guarantees; (3) an optimization‐based safe projection algorithm to realign the simplified mesh with the input, eliminating the surface shift introduced by re‐meshing and recovering the original sharp features. The algorithm demonstrates remarkable efficiency, simplifying a 2‐million‐face mesh to 20k triangles in 3 seconds on RTX4090. We evaluated the approach on the Thingi10K dataset and showcased its exceptional performance in geometry preservation and speed. https://seonghunn.github.io/pamo/ Seonghun Oh, Xiaodi Yuan, Xinyue Wei, Ruoxi Shi, Fanbo Xiang, Minghua Liu, Hao Su 0001 |
Comput. Graph. Forum | 6 |
| 2025 | Fair-DETR: Detection transformer with adaptive multi-scale attention and dual strong constraint-aware query selection
Minghua Liu, Lianen Qu, Danning Li |
Image Vis. Comput. | 2 |
| 2024 | One-2-3-45++: Fast Single Image to 3D Objects with Consistent Multi-View Generation and 3D DiffusionabstractRecent advancements in open-world 3D object generation have been remarkable, with image-to-3D methods of-fering superior fine-grained control over their text-to-3D counterparts. However, most existing models fall short in simultaneously providing rapid generation speeds and high fidelity to input images - two features essential for practi-cal applications. In this paper, we present One-2-3-45++, an innovative method that transforms a single image into a detailed 3D textured mesh in approximately one minute. Our approach aims to fully harness the extensive knowledge embedded in 2D diffusion models and priors from valuable yet limited 3D data. This is achieved by initially finetuning a 2D diffusion model for consistent multi-view image generation, followed by elevating these images to 3D with the aid of multi-view-conditioned 3D native diffusion models. Extensive experimental evaluations demonstrate that our method can produce high-quality, diverse 3D assets that closely mirror the original input image. Minghua Liu, Ruoxi Shi, Zhuoyang Zhang, Chao Xu 0016, Xinyue Wei, Hansheng Chen 0001, Chong Zeng 0001, Jiayuan Gu, Hao Su 0001 |
CVPR | 1 |
| 2024 | SpaRP: Fast 3D Object Reconstruction and Pose Estimation from Sparse Views
Chao Xu 0016, Ang Li 0010, Ruoxi Shi, Hao Su 0001, Minghua Liu |
ECCV (64) | 7 |
| 2024 | DrS: Learning Reusable Dense Rewards for Multi-Stage TasksabstractThe success of many RL techniques heavily relies on human-engineered dense rewards, which typically demands substantial domain expertise and extensive trial and error. In our work, we propose **DrS** (**D**ense **r**eward learning from **S**tages), a novel approach for learning *reusable* dense rewards for multi-stage tasks in a data-driven manner. By leveraging the stage structures of the task, DrS learns a high-quality dense reward from sparse rewards and demonstrations if given. The learned rewards can be *reused* in unseen tasks, thus reducing the human effort for reward engineering. Extensive experiments on three physical robot manipulation task families with 1000+ task variants demonstrate that our learned rewards can be reused in unseen tasks, resulting in improved performance and sample efficiency of RL algorithms. The learned rewards even achieve comparable performance to human-engineered rewards on some tasks. See our [project page](https://sites.google.com/view/iclr24drs) for more details. Tongzhou Mu, Minghua Liu, Hao Su 0001 |
ICLR | 2 |
| 2024 | MeshFormer : High-Quality Mesh Generation with 3D-Guided Reconstruction ModelabstractOpen-world 3D reconstruction models have recently garnered significant attention. However, without sufficient 3D inductive bias, existing methods typically entail expensive training costs and struggle to extract high-quality 3D meshes. In this work, we introduce MeshFormer, a sparse-view reconstruction model that explicitly leverages 3D native structure, input guidance, and training supervision. Specifically, instead of using a triplane representation, we store features in 3D sparse voxels and combine transformers with 3D convolutions to leverage an explicit 3D structure and projective bias. In addition to sparse-view RGB input, we require the network to take input and generate corresponding normal maps. The input normal maps can be predicted by 2D diffusion models, significantly aiding in the guidance and refinement of the geometry's learning. Moreover, by combining Signed Distance Function (SDF) supervision with surface rendering, we directly learn to generate high-quality meshes without the need for complex multi-stage training processes. By incorporating these explicit 3D biases, MeshFormer can be trained efficiently and deliver high-quality textured meshes with fine-grained geometric details. It can also be integrated with 2D diffusion models to enable fast single-image-to-3D and text-to-3D tasks. **Videos are available at https://meshformer3d.github.io/** Minghua Liu, Chong Zeng 0001, Xinyue Wei, Ruoxi Shi, Chao Xu 0016, Zhaoning Wang, Xiaoshuai Zhang, Isabella Liu, Hongzhi Wu, Hao Su 0001 |
NeurIPS | 1 |
| 2023 | PartSLIP: Low-Shot Part Segmentation for 3D Point Clouds via Pretrained Image-Language ModelsabstractGeneralizable 3D part segmentation is important but challenging in vision and robotics. Training deep models via conventional supervised methods requires large-scale 3D datasets with fine-grained part annotations, which are costly to collect. This paper explores an alternative way for low-shot part segmentation of 3D point clouds by leveraging a pretrained image-language model, GLIP. which achieves superior performance on open-vocabulary 2D detection. We transfer the rich knowledge from 2D to 3D through GLIP-based part detection on point cloud rendering and a novel 2D-to-3D label lifting algorithm. We also utilize multi-view 3D priors and few-shot prompt tuning to boost performance significantly. Extensive evaluation on PartNet and PartNet-Mobility datasets shows that our method enables excellent zero-shot 3D part segmentation. Our few-shot version not only outperforms existing few-shot approaches by a large margin but also achieves highly competitive results compared to the fully supervised counterpart. Furthermore, we demonstrate that our method can be directly applied to iPhone-scanned point clouds without significant domain gaps. Minghua Liu, Yinhao Zhu, Shizhong Han, Zhan Ling, Fatih Porikli, Hao Su 0001 |
CVPR | 1 |
| 2023 | Distilling Large Vision-Language Model with Out-of-Distribution GeneralizabilityabstractLarge vision-language models have achieved outstanding performance, but their size and computational requirements make their deployment on resource-constrained devices and time-sensitive tasks impractical. Model distillation, the process of creating smaller, faster models that maintain the performance of larger models, is a promising direction towards the solution. This paper investigates the distillation of visual representations in large teacher vision-language models into lightweight student models using a small- or mid-scale dataset. Notably, this study focuses on open-vocabulary out-of-distribution (OOD) generalization, a challenging problem that has been overlooked in previous model distillation literature. We propose two principles from vision and language modality perspectives to enhance student’s OOD generalization: (1) by better imitating teacher’s visual representation space, and carefully promoting better coherence in vision-language alignment with the teacher; (2) by enriching the teacher’s language representations with informative and fine-grained semantic attributes to effectively distinguish between different labels. We propose several metrics and conduct extensive experiments to investigate their techniques. The results demonstrate significant improvements in zero-shot and few-shot student performance on open-vocabulary out-of-distribution classification, highlighting the effectiveness of our proposed approaches. Code released at this link. Yunhao Fang, Minghua Liu, Zhan Ling, Zhuowen Tu, Hao Su 0001 |
ICCV | 3 |
| 2023 | OpenShape: Scaling Up 3D Shape Representation Towards Open-World UnderstandingabstractWe introduce OpenShape, a method for learning multi-modal joint representations of text, image, and point clouds. We adopt the commonly used multi-modal contrastive learning framework for representation alignment, but with a specific focus on scaling up 3D representations to enable open-world 3D shape understanding. To achieve this, we scale up training data by ensembling multiple 3D datasets and propose several strategies to automatically filter and enrich noisy text descriptions. We also explore and compare strategies for scaling 3D backbone networks and introduce a novel hard negative mining module for more efficient training. We evaluate OpenShape on zero-shot 3D classification benchmarks and demonstrate its superior capabilities for open-world recognition. Specifically, OpenShape achieves a zero-shot accuracy of 46.8% on the 1,156-category Objaverse-LVIS benchmark, compared to less than 10% for existing methods. OpenShape also achieves an accuracy of 85.3% on ModelNet40, outperforming previous zero-shot baseline methods by 20% and performing on par with some fully-supervised methods. Furthermore, we show that our learned embeddings encode a wide range of visual and semantic concepts (e.g., subcategories, color, shape, style) and facilitate fine-grained text-3D and image-3D interactions. Due to their alignment with CLIP embeddings, our learned shape representations can also be integrated with off-the-shelf CLIP-based models for various applications, such as point cloud captioning and point cloud-conditioned image generation. Minghua Liu, Ruoxi Shi, Kaiming Kuang, Yinhao Zhu, Shizhong Han, Fatih Porikli, Hao Su 0001 |
NeurIPS | 1 |
| 2023 | One-2-3-45: Any Single Image to 3D Mesh in 45 Seconds without Per-Shape OptimizationabstractSingle image 3D reconstruction is an important but challenging task that requires extensive knowledge of our natural world. Many existing methods solve this problem by optimizing a neural radiance field under the guidance of 2D diffusion models but suffer from lengthy optimization time, 3D inconsistency results, and poor geometry. In this work, we propose a novel method that takes a single image of any object as input and generates a full 360-degree 3D textured mesh in a single feed-forward pass. Given a single image, we first use a view-conditioned 2D diffusion model, Zero123, to generate multi-view images for the input view, and then aim to lift them up to 3D space. Since traditional reconstruction methods struggle with inconsistent multi-view predictions, we build our 3D reconstruction module upon an SDF-based generalizable neural surface reconstruction method and propose several critical training strategies to enable the reconstruction of 360-degree meshes. Without costly optimizations, our method reconstructs 3D shapes in significantly less time than existing methods. Moreover, our method favors better geometry, generates more 3D consistent results, and adheres more closely to the input image. We evaluate our approach on both synthetic data and in-the-wild images and demonstrate its superiority in terms of both mesh quality and runtime. In addition, our approach can seamlessly support the text-to-3D task by integrating with off-the-shelf text-to-image diffusion models. Minghua Liu, Chao Xu 0016, Haian Jin, Mukund Varma T., Zexiang Xu, Hao Su 0001 |
NeurIPS | 1 |
| 2023 | Analysis and test of influence of memristor non-ideal characteristics on facial expression recognition accuracy
Yening Li, Ruoyu Meng, Minghua Liu |
Expert Syst. Appl. | 5 |
| 2023 | Close the Optical Sensing Domain Gap by Physics-Grounded Active Stereo Sensor SimulationabstractIn this article, we focus on the simulation of active stereovision depth sensors, which are popular in both academic and industry communities. Inspired by the underlying mechanism of the sensors, we designed a fully physics-grounded simulation pipeline that includes material acquisition, ray-tracing-based infrared (IR) image rendering, IR noise simulation, and depth estimation. The pipeline is able to generate depth maps with material-dependent error patterns similar to a real depth sensor in real time. We conduct real experiments to show that perception algorithms and reinforcement learning policies trained in our simulation platform could transfer well to the real-world test cases without any fine-tuning. Furthermore, due to the high degree of realism of this simulation, our depth sensor simulator can be used as a convenient testbed to evaluate the algorithm performance in the real world, which will largely reduce the human effort in developing robotic algorithms. The entire pipeline has been integrated into the SAPIEN simulator and is open-sourced to promote the research of vision and robotics communities. Xiaoshuai Zhang, Rui Chen 0019, Ang Li 0010, Fanbo Xiang, Yuzhe Qin, Jiayuan Gu, Zhan Ling, Minghua Liu, Peiyu Zeng, Songfang Han, Zhiao Huang, Tongzhou Mu, Jing Xu 0011, Hao Su 0001 |
IEEE Trans. Robotics | 8 |
| 2022 | LESS: Label-Efficient Semantic Segmentation for LiDAR Point Clouds
Minghua Liu, Charles R. Qi, Boqing Gong, Hao Su 0001, Dragomir Anguelov |
ECCV (39) | 1 |
| 2022 | Approximate convex decomposition for 3D meshes with collision-aware concavity and tree searchabstractApproximate convex decomposition aims to decompose a 3D shape into a set of almost convex components, whose convex hulls can then be used to represent the input shape. It thus enables efficient geometry processing algorithms specifically designed for convex shapes and has been widely used in game engines, physics simulations, and animation. While prior works can capture the global structure of input shapes, they may fail to preserve fine-grained details (e.g., filling a toaster's slots), which are critical for retaining the functionality of objects in interactive environments. In this paper, we propose a novel method that addresses the limitations of existing approaches from three perspectives: (a) We introduce a novel collision-aware concavity metric that examines the distance between a shape and its convex hull from both the boundary and the interior. The proposed concavity preserves collision conditions and is more robust to detect various approximation errors. (b) We decompose shapes by directly cutting meshes with 3D planes. It ensures generated convex hulls are intersection-free and avoids voxelization errors. (c) Instead of using a one-step greedy strategy, we propose employing a multi-step tree search to determine the cutting planes, which leads to a globally better solution and avoids unnecessary cuttings. Through extensive evaluation on a large-scale articulated object dataset, we show that our method generates decompositions closer to the original shape with fewer components. It thus supports delicate and efficient object interaction in downstream applications. Xinyue Wei, Minghua Liu, Zhan Ling, Hao Su 0001 |
ACM Trans. Graph. | 2 |
| 2021 | DeepMetaHandles: Learning Deformation Meta-Handles of 3D Meshes With Biharmonic CoordinatesabstractWe propose DeepMetaHandles, a 3D conditional generative model based on mesh deformation. Given a collection of 3D meshes of a category and their deformation handles (control points), our method learns a set of meta-handles for each shape, which are represented as combinations of the given handles. The disentangled meta-handles factorize all the plausible deformations of the shape, while each of them corresponds to an intuitive deformation. A new deformation can then be generated by sampling the co-efficients of the meta-handles in a specific range. We employ biharmonic coordinates as the deformation function, which can smoothly propagate the control points’ translations to the entire mesh. To avoid learning zero deformaion as meta-handles, we incorporate a target-fitting module which deforms the input mesh to match a random target. To enhance deformations’ plausibility, we employ a soft-rasterizer-based discriminator that projects the meshes to a 2D space. Our experiments demonstrate the superiority of the generated deformations as well as the interpretability and consistency of the learned meta-handles. The code is available at https://github.com/Colin97/DeepMetaHandles. Minghua Liu, Minhyuk Sung, Radomír Mech, Hao Su 0001 |
CVPR | 1 |
| 2020 | Morphing and Sampling Network for Dense Point Cloud Completionabstract3D point cloud completion, the task of inferring the complete geometric shape from a partial point cloud, has been attracting attention in the community. For acquiring high-fidelity dense point clouds and avoiding uneven distribution, blurred details, or structural loss of existing methods' results, we propose a novel approach to complete the partial point cloud in two stages. Specifically, in the first stage, the approach predicts a complete but coarse-grained point cloud with a collection of parametric surface elements. Then, in the second stage, it merges the coarse-grained prediction with the input point cloud by a novel sampling algorithm. Our method utilizes a joint loss function to guide the distribution of the points. Extensive experiments verify the effectiveness of our method and demonstrate that it outperforms the existing methods in both the Earth Mover's Distance (EMD) and the Chamfer Distance (CD). Minghua Liu, Lu Sheng, Sheng Yang 0007, Shi-Min Hu 0001 |
AAAI | 1 |
| 2020 | SAPIEN: A SimulAted Part-Based Interactive ENvironmentabstractBuilding home assistant robots has long been a goal for vision and robotics researchers. To achieve this task, a simulated environment with physically realistic simulation, sufficient articulated objects, and transferability to the real robot is indispensable. Existing environments achieve these requirements for robotics simulation with different levels of simplification and focus. We take one step further in constructing an environment that supports household tasks for training robot learning algorithm. Our work, SAPIEN, is a realistic and physics-rich simulated environment that hosts a large-scale set of articulated objects. SAPIEN enables various robotic vision and interaction tasks that require detailed part-level understanding.We evaluate state-of-the-art vision algorithms for part detection and motion attribute recognition as well as demonstrate robotic interaction tasks using heuristic approaches and reinforcement learning algorithms. We hope that SAPIEN will open research directions yet to be explored, including learning cognition through interaction, part motion discovery, and construction of robotics-ready simulated game environment. Fanbo Xiang, Yuzhe Qin, Kaichun Mo, Yikuan Xia, Hao Zhu 0008, Fangchen Liu, Minghua Liu, Hanxiao Jiang 0001, Yifu Yuan, He Wang 0010, Li Yi 0001, Angel X. Chang, Leonidas J. Guibas, Hao Su 0001 |
CVPR | 7 |
| 2020 | Meshing Point Clouds with Predicted Intrinsic-Extrinsic Ratio Guidance
Minghua Liu, Xiaoshuai Zhang, Hao Su 0001 |
ECCV (8) | 1 |
| 2020 | Multi-task Batch Reinforcement Learning with Metric LearningabstractWe tackle the Multi-task Batch Reinforcement Learning problem. Given multiple datasets collected from different tasks, we train a multi-task policy to perform well in unseen tasks sampled from the same distribution. The task identities of the unseen tasks are not provided. To perform well, the policy must infer the task identity from collected transitions by modelling its dependency on states, actions and rewards. Because the different datasets may have state-action distributions with large divergence, the task inference module can learn to ignore the rewards and spuriously correlate \textit{only} state-action pairs to the task identity, leading to poor test time performance. To robustify task inference, we propose a novel application of the triplet loss. To mine hard negative examples, we relabel the transitions from the training tasks by approximating their reward functions. When we allow further training on the unseen tasks, using the trained policy as an initialization leads to significantly faster convergence compared to randomly initialized policies (up to 80% improvement and across 5 different Mujoco task distributions). We name our method \textbf{MBML} (\textbf{M}ulti-task \textbf{B}atch RL with \textbf{M}etric \textbf{L}earning). Minghua Liu, Kamil Ciosek, Henrik I. Christensen, Hao Su 0001 |
NeurIPS | 4 |
| 2020 | HeteroFusion: Dense Scene Reconstruction Integrating Multi-SensorsabstractWe present a novel approach to integrate data from multiple sensor types for dense 3D reconstruction of indoor scenes in realtime. Existing algorithms are mainly based on a single RGBD camera and thus require continuous scanning of areas with sufficient geometric features. Otherwise, tracking may fail due to unreliable frame registration. Inspired by the fact that the fusion of multiple sensors can combine their strengths towards a more robust and accurate self-localization, we incorporate multiple types of sensors which are prevalent in modern robot systems, including a 2D range sensor, an inertial measurement unit (IMU), and wheel encoders. We fuse their measurements to reinforce the tracking process and to eventually obtain better 3D reconstructions. Specifically, we develop a 2D truncated signed distance field (TSDF) volume representation for the integration and ray-casting of laser frames, leading to a unified cost function in the pose estimation stage. For validation of the estimated poses in the loop-closure optimization process, we train a classifier for the features extracted from heterogeneous sensors during the registration progress. To evaluate our method on challenging use case scenarios, we assembled a scanning platform prototype to acquire real-world scans. We further simulated synthetic scans based on high-fidelity synthetic scenes for quantitative evaluation. Extensive experimental evaluation on these two types of scans demonstrate that our system is capable of robustly acquiring dense 3D reconstructions and outperforms state-of-the-art RGBD and LiDAR systems. Sheng Yang 0007, Beichen Li 0005, Minghua Liu, Yukun Lai, Leif Kobbelt, Shi-Min Hu 0001 |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2018 | A path detecting method to analyze the interactive compatibility of service processes based on WS-BPELabstractSummary Petri nets are frequently used formal tools to analyze the compatibility of interactive service processes described by Web Services Business Process Execution Language (WS‐BPEL). However, the traditional methods based on Petri nets were with a high computable complexity for state space explosion. To resolve such problem, a logic Petri net–based path detecting method for compatibility analysis of interactive service processes is proposed. From the provided mapping rules, the service process described by WS‐BPEL is modeled as a service net based on logic Petri nets. The evaluation of interactive compatibility of two service processes is converted to analyze whether their service nets can be composed as a non‐blocked synthetic service net. The non‐blocked property is checked by detecting the reachability of the potential connected paths in a service net. To reduce the complexity of computing the connected paths in a service net, we propose a merge‐reduced method to generate the path expression of its skeleton service net. The potential connected paths of a service net can be obtained by unfolding the path expression. Compared with the traditional method based Petri nets, the proposed method is with high efficiency and it can greatly alleviate the problem of state space explosion in analyzing interactive compatibility of service processes. Qiang Hu 0002, Minghua Liu, Zhen Zhao 0006, Junwei Du |
Concurr. Comput. Pract. Exp. | 2 |
| 2017 | Saliency-aware Real-time Volumetric Fusion for Object ReconstructionabstractAbstract We present a real‐time approach for acquiring 3D objects with high fidelity using hand‐held consumer‐level RGB‐D scanning devices. Existing real‐time reconstruction methods typically do not take the point of interest into account, and thus might fail to produce clean reconstruction results of desired objects due to distracting objects or backgrounds. In addition, any changes in background during scanning, which can often occur in real scenarios, can easily break up the whole reconstruction process. To address these issues, we incorporate visual saliency into a traditional real‐time volumetric fusion pipeline. Salient regions detected from RGB‐D frames suggest user‐intended objects, and by understanding user intentions our approach can put more emphasis on important targets, and meanwhile, eliminate disturbance of non‐important objects. Experimental results on real‐world scans demonstrate that our system is capable of effectively acquiring geometric information of salient objects in cluttered real‐world scenes, even if the backgrounds are changing. Sheng Yang 0007, Minghua Liu, Hongbo Fu 0001, Shi-Min Hu 0001 |
Comput. Graph. Forum | 3 |
| 2014 | A Multiagent Q-Learning-Based Optimal Allocation Approach for Urban Water Resource Management SystemabstractWater environment system is a complex system, and an agent-based model presents an effective approach that has been implemented in water resource management research. Urban water resource optimal allocation is a challenging and critical issue in water environment systems, which belongs to the resource optimal allocation problem. In this paper, a novel approach based on multiagent Q-learning is proposed to deal with this problem. In the proposed approach, water users of different regions in the city are abstracted into the agent-based model. To realize the cooperation among these stakeholder agents, a maximum mapping value function-based Q-learning algorithm is proposed in this study, which allows the agents to self-learn. In the proposed algorithm, an adaptive reward value function is used to improve the performance of the multiagent Q-learning algorithm, where the influence of multiple factors on the optimal allocation can be fully considered. The proposed approach can deal with various situations in urban water resource allocation. The experimental results show that the proposed approach is capable of allocating water resource efficiently and the objectives of all the stakeholder agents can be successfully achieved. Jianjun Ni, Minghua Liu, Simon X. Yang |
IEEE Trans Autom. Sci. Eng. | 2 |