Jianheng Liu

dblp:309/2328 · DBLP profile ↗
← Back
9ranked-venue papers
3as first author
9since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 3 first-author · 9 since 2021Systems, architecture and hardware · 7 · 3 first-author · 7 since 2021
YearPublicationVenuePosition
2025 DistRL: An Asynchronous Distributed Reinforcement Learning Framework for On-Device Control Agent
abstract
On-device control agents, especially on mobile devices, are responsible for operating mobile devices to fulfill users' requests, enabling seamless and intuitive interactions. Integrating Multimodal Large Language Models (MLLMs) into these agents enhances their ability to understand and execute complex commands, thereby improving user experience. However, fine-tuning MLLMs for on-device control presents significant challenges due to limited data availability and inefficient online training processes. This paper introduces DistRL, a novel framework designed to enhance the efficiency of online RL fine-tuning for mobile device control agents. DistRL employs centralized training and decentralized data acquisition to ensure efficient fine-tuning in the context of dynamic online interactions. Additionally, the framework is backed by our tailor-made RL algorithm, which effectively balances exploration with the prioritized utilization of collected data to ensure stable and robust training. Our experiments show that, on average, DistRL delivers a 3$\times$ improvement in training efficiency and enables training data collection 2.4$\times$ faster than the leading synchronous multi-machine methods. Notably, after training, DistRL achieves a 20\% relative improvement in success rate compared to state-of-the-art methods on general Android tasks from an open benchmark, significantly outperforming existing approaches while maintaining the same training time. These results validate DistRL as a scalable and efficient solution, offering substantial improvements in both training efficiency and agent performance for real-world, in-the-wild device control tasks.
Taiyi Wang, Jianheng Liu, Jianye Hao, Jun Wang 0012, Kun Shao
ICLR3
2025 Neural Surface Reconstruction and Rendering for LiDAR-Visual Systems
abstract
This paper presents a unified surface reconstruction and rendering framework for LiDAR-visual systems, integrating Neural Radiance Fields (NeRF) and Neural Distance Fields (NDF) to recover both appearance and structural information from posed images and point clouds. We address the structural visible gap between NeRF and NDF by utilizing a visible-aware occupancy map to classify space into the free, occupied, visible unknown, and background regions. This classification facilitates the recovery of a complete appearance and structure of the scene. We unify the training of the NDF and NeRF using a spatial-varying scale SDF-to-density transformation for levels of detail for both structure and appearance. The proposed method leverages the learned NDF for structure-aware NeRF training by an adaptive sphere tracing sampling strategy for accurate structure rendering. In return, NeRF further refines structural in recovering missing or fuzzy structures in the NDF. Extensive experiments demonstrate the superior quality and versatility of the proposed method across various scenarios. To benefit the community, the codes will be released at https://github.com/hku-mars/M2Mapping.
Jianheng Liu, Chunran Zheng, Yunfei Wan, Yixi Cai, Fu Zhang 0002
ICRA1
2025 OCMDP: Observation-Constrained Markov Decision Process
abstract
In many practical applications, decision-making processes must balance the costs of acquiring information with the benefits it provides. Traditional control systems often assume full observability, an unrealistic assumption when observations are expensive. We tackle the challenge of simultaneously learning observation and control strategies in such cost-sensitive environments by introducing the Observation-Constrained Markov Decision Process (OCMDP), where the policy influences the observability of the true state. To manage the complexity arising from the combined observation and control actions, we develop an iterative, model-free deep reinforcement learning algorithm that separates the sensing and control components of the policy. This decomposition enables efficient learning in the expanded action space by focusing on when and what to observe, as well as determining optimal control actions, without requiring knowledge of the environment’s dynamics. Experimental results across diverse healthcare tasks and environments demonstrate that our proposed approach substantially reduces observation costs while significantly outperforming baseline methods in efficiency across various real-world scenarios.
Taiyi Wang, Jianheng Liu, Bryan Lee
IJCNN2
2025 Efficient Swept Volume-Based Trajectory Generation for Arbitrary-Shaped Ground Robot Navigation
abstract
Navigating an arbitrary-shaped ground robot safely in cluttered environments remains a challenging problem. The existing trajectory planners that account for the robot’s physical geometry severely suffer from the intractable runtime. To achieve both computational efficiency and Continuous Collision Avoidance (CCA) of arbitrary-shaped ground robot planning, we proposed a novel coarse-to-fine navigation framework that significantly accelerates planning. In the first stage, a sampling-based method selectively generates distinct topological paths that guarantee a minimum inflated margin. In the second stage, a geometry-aware front-end strategy is designed to discretize these topologies into full-state robot motion sequences while concurrently partitioning the paths into SE(2) sub-problems and simpler ℝ2sub-problems for back-end optimization. In the final stage, an SVSDF-based optimizer generates trajectories tailored to these sub-problems and seamlessly splices them into a continuous final motion plan. Extensive benchmark comparisons show that the proposed method is one to several orders of magnitude faster than the cutting-edge methods in runtime while maintaining a high planning success rate and ensuring CCA.
Yisheng Li, Longji Yin, Yixi Cai, Jianheng Liu, Fangcheng Zhu, Mingpu Ma, Siqi Liang 0004, Fu Zhang 0002
IROS4
2025 GS-SDF: LiDAR-Augmented Gaussian Splatting and Neural SDF for Geometrically Consistent Rendering and Reconstruction
abstract
Digital twins are fundamental to the development of autonomous driving and embodied artificial intelligence. However, achieving high-granularity surface reconstruction and high-fidelity rendering remains a challenge. Gaussian splatting offers efficient photorealistic rendering but struggles with geometric inconsistencies due to fragmented primitives and sparse observational data in robotics applications. Existing regularization methods, which rely on render-derived constraints, often fail in complex environments. Moreover, effectively integrating sparse LiDAR data with Gaussian splatting remains challenging. We propose a unified LiDAR-visual system that synergizes Gaussian splatting with a neural signed distance field. The accurate LiDAR point clouds enable a trained neural signed distance field to offer a manifold geometry field. This motivates us to offer an SDF-based Gaussian initialization for physically grounded primitive placement and a comprehensive geometric regularization for geometrically consistent rendering and reconstruction. Experiments demonstrate superior reconstruction accuracy and rendering quality across diverse trajectories. To benefit the community, the codes are released at https: //github.com/hku-mars/GS-SDF.
Jianheng Liu, Yunfei Wan, Chunran Zheng, Jiarong Lin, Fu Zhang 0002
IROS1
2025 Mesh-Learner: Texturing Mesh with Spherical Harmonics
abstract
In this paper, we present a 3D reconstruction and rendering framework termed Mesh-Learner that is natively compatible with traditional rasterization pipelines. It integrates mesh and spherical harmonic (SH) Texture (i.e., texture filled with SH coefficients) into the learning process to learn each mesh’s view-dependent radiance end-to-end. Images are rendered by interpolating surrounding SH Texels at each pixel’s sampling point using a novel interpolation method. Conversely, gradients from each pixel are back-propagated to the related SH Texels in SH Textures. Mesh-Learner exploits graphic features of rasterization pipeline (texture sampling, deferred rendering) to render, which makes Mesh-Learner naturally compatible with tools (e.g., Blender) and tasks (e.g., 3D reconstruction, scene rendering, reinforcement learning for robotics) that are based on rasterization pipelines. Our system can train vast, unlimited scenes because we transfer only the SH Textures within the frustum to the GPU for training. At other times, the SH Textures are stored in CPU RAM, which results in moderate GPU memory usage. The rendering results on interpolation and extrapolation sequences in the Replica and FAST-LIVO2 datasets achieve state-of-the-art performance compared to existing state-of-the-art methods (e.g., 3D Gaussian Splatting and M2-Mapping). To benefit the society, the code will be available at https://github.com/hku-mars/Mesh-Learner.
Yunfei Wan, Jianheng Liu, Chunran Zheng, Jiarong Lin, Fu Zhang 0002
IROS2
2024 Towards Large-Scale Incremental Dense Mapping using Robot-centric Implicit Neural Representation
abstract
Large-scale dense mapping is vital in robotics, digital twins, and virtual reality. Recently, implicit neural mapping has shown remarkable reconstruction quality. However, incremental large-scale mapping with implicit neural representations remains problematic due to low efficiency, limited video memory, and the catastrophic forgetting phenomenon. To counter these challenges, we introduce the Robot-centric Implicit Mapping (RIM) technique for large-scale incremental dense mapping. This method employs a hybrid representation, encoding shapes with implicit features via a multi-resolution voxel map and decoding signed distance fields through a shallow MLP. We advocate for a robot-centric local map to boost model training efficiency and curb the catastrophic forgetting issue. A decoupled scalable global map is further developed to archive learned features for reuse and maintain constant video memory consumption. Validation experiments demonstrate our method’s exceptional quality, efficiency, and adaptability across diverse scales and scenes over advanced dense mapping methods using range sensors. Our system’s code will be accessible at https://github.com/HITSZ-NRSL/RIM.git.
Jianheng Liu, Haoyao Chen
ICRA1
2022 Sampling-Based View Planning for MAVs in Active Visual-inertial State Estimation
abstract
Micro aerial vehicles usually have strap-down sensors on the vehicle body, leading to the severe coupling effect between perception and trajectory planning. As a result, visual-inertial simultaneous localization and mapping (VI-SLAM) technologies implemented on MAVs suffer from tracking failure problems, especially in featureless environments. To overcome these challenges, based on MAVs with movable camera mechanisms (e.g., gimbal stabilizer, pan-tilt, or bionic neck-eye system), we proposed two sampling-based algorithms for known and unknown environments respectively. The first active perception planning algorithm based on a scene richness model is developed with a built feature map for the environment. Differ from the first algorithm, the second one is modified for active localization in unknown 3D space. It is basically a time-based sampling-based approach that uses the same scene richness model. In addition, it also achieved a balance between exploitation and exploration. With the above solutions, the robustness of visual perception is improved while avoiding over-exploitation of known information. Simulation and real-world experiments are performed to verify the feasibility of our algorithms.
Zhengyu Hua, Fengyu Quan, Haoyao Chen, Jiabi Sun, Jianheng Liu, Yun-Hui Liu 0001
IROS5
2021 Vision-encoder-based Payload State Estimation for Autonomous MAV With a Suspended Payload
abstract
Autonomous delivery of suspended payloads with MAVs has many applications in rescue and logistics transportation. Robust and online estimation of the payload status is important but challenging especially in outdoor environments. The paper develops a novel real-time system for estimating the payload position; the system consists of a monocular fisheye camera and a novel encoder-based device. A Gaussian fusion-based estimation algorithm is developed to obtain the payload state estimation. Based on the robust payload position estimation, a payload controller is presented to ensure the re-liable tracking performance on aggressive trajectories. Several experiments are performed to validate the high performance of the proposed method.
Yunfan Ren, Jianheng Liu, Haoyao Chen, Yun-Hui Liu 0001
IROS2