Kaihong Huang

dblp:130/8023 · DBLP profile ↗
← Back
14ranked-venue papers
5as first author
8since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 4 first-author · 6 since 2021Systems, architecture and hardware · 9 · 5 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021
YearPublicationVenuePosition
2026 Compressing LLM Knowledge into Graph Representations for Text-attributed Graphs Learning
abstract
Text-attributed graphs (TAGs) require jointly modeling relational structure and node-level text.Existing GNN-LLM approaches perform by incorporating large language models at inference time for processing the text attributes, resulting in costly deployment.More fundamentally, LLM knowledge is typically used in a sample-wise manner, leading to inefficient utilization across graph instances.In this work, we study how interactions with LLM embedding spaces affect graph representations, and show that projecting into the LLM space can learn better GNNs.That is to say, the knowledge encoded in LLM embeddings can be compressed into graph representations.Based on this insight, we propose a framework that internalizes LLM knowledge within graph models and supports inference-efficient TAG learning.Our framework employs a hierarchical Proxy-Purifier module with distribution-level regularization, using LLM embeddings only as training-time guidance.With this module, the model operates TAGs without invoking LLMs, achieving high efficiency as standard GNNs without LLMs.Notably, experiments on five popular TAG tasks further demonstrate that our method can also achieve consistent performance gains, in comparison to existing GNN-LLM approaches.
Runhuai Chen, Dian Shen, Kaihong Huang, Beilun Wang
ACL (1)4
2026 Lifelong memory organization: Incremental multi-modal data fusion via Bayesian clustering and large language models
Junfeng Shi, Hainan Pan, Kaihong Huang
Pattern Recognit. Lett.3
2026 Active Learning-Based Joint Optimization of Bionic Nose Profile and Water-Entry Strategy for Aerial-Aquatic Robots
abstract
The practical application of aerial-aquatic robots is hindered by severe impact loads during water entry. Existing load reduction methods are inefficient for robots requiring rapid, high-frequency aerial-aquatic transitions. Therefore, this study proposes an active learning based joint optimization framework for nose profiles and water-entry strategies, which is fully automated. Specifically, an 8-dimensional parameter space is defined for generating dataset inputs through Latin Hypercube Sampling (LHS). Furthermore, a Deep Kernel Learning (DKL) surrogate model is trained under different water-entry strategies, in order to predict peak impact loads and quantify prediction uncertainty for varying nose profiles. Within each active learning loop, the Non-dominated Sorting Genetic Algorithm III (NSGA-III) guides the selective labelling of samples to expand the dataset, and the DKL model is iteratively retrained until convergence. Compared against state-of-the-art methods, the proposed approach reduces the number of required high-fidelity Computational Fluid Dynamics (CFD) simulations to 65% of that of the comparison methods, while achieving maximum reductions of 66.9% in axial and 72.8% in normal impact loads across the parameter space, respectively. Notably, under the water-entry strategies employed by the hunting behavior of gannets, the approach yields profiles closely matching the skull morphology of the northern gannet. In this case, this work not only delivers an efficient and reliable impact loads reduction solution for aerial-aquatic robots, but also reveals the role of natural selection in minimizing such loads.
Mengsen Zhao, Kaihong Huang, Shiyou Zhao, Huimin Lu 0002, Junhao Xiao 0001
IEEE Trans Autom. Sci. Eng.2
2025 Efficient Instance Motion-Aware Point Cloud Scene Prediction
abstract
Point cloud prediction (PCP) aims to forecast future 3D point clouds of scenes by leveraging sequential historical LiDAR scans, offering a promising avenue to enhance the perceptual capabilities of autonomous systems. However, existing methods mostly adopt an end-to-end approach without explicitly modeling moving instances, limiting their effectiveness in dynamic real-world environments. In this paper, we propose IMPNet, a novel instance motion-aware network for future point cloud scene prediction. Unlike prior works, IMPNet explicitly incorporates motion and instance-level information to enhance PCP accuracy. Specifically, we extract appearance and motion features from range images and residual images using a dual-branch convolutional network and fuse them via a motion attention block. Our framework further integrates a motion head for identifying moving objects and an instance-assisted training strategy to improve instance-wise point cloud predictions. Extensive experiments on multiple datasets demonstrate that our proposed network achieves state-of-the-art (SOTA) performance in PCP with superior predictive accuracy and robust generalization across diverse driving scenarios. Our method has been released at https://github.com/nubot-nudt/IMPNet.
Xieyuanli Chen, Kaihong Huang, Huimin Lu 0002
IROS4
2025 C-TRAC: Terrain-Adaptive Control for Articulated Tracked Robots via Contact-Aware Reinforcement Learning
abstract
Articulated tracked robots face significant challenges in maintaining stable locomotion over uneven terrain due to unknown contact points between tracks and ground, which are critical for dynamic control. Unlike legged robots, where contact locations can be predicted, tracked systems require real-time adaptation to varying terrains. This paper presents C-TRAC, a terrain-adaptive control framework that integrates reinforcement learning with a contact-modeling variational autoencoder (C-VAE) to enable robust obstacle traversal. We first train a C-VAE in simulation to reconstruct high-fidelity contact information (position and binary probability) from noisy sensor measurements. This model learns a latent representation of terrain contacts, capturing complex interactions between the robot’s kinematics and environment. Subsequently, we employ an asymmetric Soft Actor-Critic (SAC) algorithm to optimize a control policy that leverages the predicted contact data for adaptive track control during locomotion. Extensive experiments validate C-TRAC in both simulated and real-world scenarios. In benchmark tests against state-of-the-art (SOTA) methods using RoboCup Rescue Robot League environments, our approach achieves superior obstacle traversal speed (up to 66.67% faster on 45◦staircase) and stability (up to 47.53% more stable on the oblique terrace) compared to contact-agnostic RL baselines and model-based methods. Notably, zero-shot sim-to-real transfer demonstrates consistent performance in unstructured outdoor ruins, also confirming the framework’s practicality.
Hainan Pan, Kaihong Huang, Xieyuanli Chen, Hongchuan Zhang, Junfeng Shi, Chuang Cheng, Bailiang Chen, Huimin Lu 0002
IROS2
2025 Autonomous Subtask Generation for Indoor Search and Rescue Mission via Large-Language-Model and Behavior-Tree Integration
abstract
The ability of autonomous subtask generation is important for robots to effectively cope with unforeseen situations during indoor search and rescue missions. While prior work mainly focused on improving individual low-level skills of the rescue robot, this paper proposes AutoExpand: a high-level framework that takes advantage of the extensive knowledge and reasoning abilities inherent in large language models (LLM) to understand human instructions and environmental situation. Through tight coupling LLM with behavior tree, our method enables the robot to autonomously generate reactive context-aware operational subtasks on-site without human intervention or additional training. A series of real-world experiments demonstrate that AutoExpand can effectively generate appropriate tasks for search and rescue missions, leading to a search scope increased by 34.45% when compared with traditional methods. The sample code is available at https://github.com/nubot-nudt/AutoExpand.
Junfeng Shi, Kaihong Huang, Hainan Pan, Junpeng Xu, Chuang Cheng, Hui Zhang 0053
IROS2
2024 Hynify: A High-throughput and Unified Accelerator for Multi-Mode Nonparametric Statistics
abstract
Nonparametric statistics methods are a class of robust and potent machine learning operators, which are widely used in various domains such as finance, medicine, and computer science. Such methods deliver an accurate estimation without an assumed data distribution. Moreover, they can handle discrete data with various data sources. Despite their desirable features, the calculation of large-scale nonparametric statistics is both compute- and memory-intensive, and the performance overhead hinders them from widespread usage.
Kaihong Huang, Dian Shen, Juntao Yang, Beilun Wang
DAC1
2023 ElC-OIS: Ellipsoidal Clustering for Open-World Instance Segmentation on LiDAR Data
abstract
Open-world Instance Segmentation (OIS) is a challenging task that aims to accurately segment every object instance appearing in the current observation, regardless of whether these instances have been labeled in the training set. This is important for safety-critical applications such as robust autonomous navigation. In this paper, we present a flexible and effective OIS framework for LiDAR point cloud that can accurately segment both known and unknown instances (i.e., seen and unseen instance categories during training). It first identifies points belonging to known classes and removes the back-ground by leveraging close-set panoptic segmentation networks. Then, we propose a novel ellipsoidal clustering method that is more adapted to the characteristic of LiDAR scans and allows precise segmentation of unknown instances. Furthermore, a diffuse searching method is proposed to handle the common over-segmentation problem presented in the known instances. With the combination of these techniques, we are able to achieve accurate segmentation for both known and unknown instances. We evaluated our method on the SemanticKITTI open-world LiDAR instance segmentation dataset. The experimental results suggest that it outperforms current state-of-the-art methods, especially with a 10.0% improvement in association quality. The source code of our method will be publicly available at https://github.com/nubot-nudt/ElC-OIS.
Wenbang Deng, Kaihong Huang, Qinghua Yu, Huimin Lu 0002, Zhiqiang Zheng 0002, Xieyuanli Chen
IROS2
2020 A Real-Time Sliding-Window-Based Visual-Inertial Odometry for MAVs
abstract
This article presents a sliding widow-based visual-inertial odometry to deal with the micro air vehicle (MAV) pose estimation problem. Errors caused by inertial measurement unit (IMU) preintegration, visual landmarks reprojection, and marginalization, are unified into a nonlinear residual minimization framework. Furthermore, a dual-step marginalization method has been proposed to increase the computational efficiency. Experiments have been conducted on publicly available datasets, as well as customized handheld and MAV platform, where state-of-the-art approaches have served as the baselines for comparison. According to the results, the proposed method has a comparative accuracy, which can run in real-time on an onboard minicomputer.
Junhao Xiao 0001, Dan Xiong, Qinghua Yu, Kaihong Huang, Huimin Lu 0002
IEEE Trans. Ind. Informatics4
2019 Accurate Direct Visual-Laser Odometry with Explicit Occlusion Handling and Plane Detection
abstract
In this paper, we address the problem of combining 3D laser scanner and camera information to estimate the motion of a mobile platform. We propose a direct laser-visual odometry approach building upon photometric image alignment. Our approach is designed to maximize the information usage of both, the image and the laser scan, to compute an accurate frame-to-frame motion estimate. To deal with the sparsity of the range measurements, our approach identifies planar point sets within individual point clouds and subsequently extract their corresponding pixel patches from the camera image. The extracted planar image patches are used together with the non-planar pixels to estimate the frame-to-frame motion using a homography formulation capable of incorporating both types of pixel alignments. To achieve high estimation accuracy, we explicitly predict possible occlusions caused by observations taken from different locations. We evaluate our proposed approach using the KITTI dataset as well as data recorded with a Clearpath Husky platform. The experiments suggest that our approach can achieve competitive estimation accuracy and produce consistently registered, colored point clouds.
Kaihong Huang, Junhao Xiao 0001, Cyrill Stachniss
ICRA1
2018 On Geometric Models and Their Accuracy for Extrinsic Sensor Calibration
abstract
Extrinsic sensor calibration is an important task in robotics. There are various ways to perform the calibration task, but it often remains unclear which methods are better than the others. In this paper, we provide a systematic study about the calibration accuracy of three types of calibration methods, each represented by an abstract geometric model based on the sensor configuration and the calibration setup. We discuss the advantages and disadvantages of each model and perform a rigorous study on their noise sensitivity from a geometric perspective. As a result, we can reveal and quantify the relative calibration accuracies of the three models, thus answering the question of “which model is better and why?”. Beside our analytical analysis, we also provide numerical simulation experiments that validate our findings.
Kaihong Huang, Cyrill Stachniss
ICRA1
2018 Joint Ego-motion Estimation Using a Laser Scanner and a Monocular Camera Through Relative Orientation Estimation and 1-DoF ICP
abstract
Pose estimation and mapping are key capabilities of most autonomous vehicles and thus a number of localization and SLAM algorithms have been developed in the past. Autonomous robots and cars are typically equipped with multiple sensors. Often, the sensor suite includes a camera and a laser range finder. In this paper, we consider the problem of incremental ego-motion estimation, using both, a monocular camera and a laser range finder jointly. We propose a new algorithm, that exploits the advantages of both sensors-the ability of cameras to determine orientations well and the ability of laser range finders to estimate the scale and to directly obtain 3D point clouds. Our approach estimates the 5 degrees of freedom relative orientation from image pairs through feature point correspondences and formulates the remaining scale estimation as a new variant of the iterative closest point problem with only one degree of freedom. We furthermore exploit the camera information in a new way to constrain the data association between laser point clouds. The experiments presented in this paper suggest that our approach is able to accurately estimate the ego-motion of a vehicle and that we obtain more accurate frame-to-frame alignments than with one sensor modality alone.
Kaihong Huang, Cyrill Stachniss
IROS1
2017 Extrinsic multi-sensor calibration for mobile robots using the Gauss-Helmert model
abstract
Most state estimation procedures in mobile robotics require information about the locations of the individual sensors on the platform. In this paper, we study the motion-based multi-sensor extrinsic calibration problem and point out an overlooked defect of traditional least squares estimation in this context. We present a novel calibration approach based-on the Gauss-Helmert estimation paradigm, together with a formulation for multi-sensor motion constraint. Our approach estimates not only the extrinsic parameters but also the pose observation errors, thus recovering the underlying sensor movements that exactly fulfill the motion constraints. Compared to traditional least squares approaches that estimate only the parameters, our approach is statistically optimal, thus is more accurate and robust. We implemented our approach and tested it on real robot. The experiments show that our approach is able to accurately determine the extrinsic configuration of each sensor and can largely improve the accuracy when the noise level is high.
Kaihong Huang, Cyrill Stachniss
IROS1
2012 A Robust Place Recognition Algorithm Based on Omnidirectional Vision for Mobile Robots
Huimin Lu 0002, Kaihong Huang, Dan Xiong, Zhiqiang Zheng 0002
RoboCup2