Yiheng Han

dblp:202/4861 · DBLP profile ↗
← Back
15ranked-venue papers
5as first author
11since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 3 first-author · 7 since 2021Systems, architecture and hardware · 5 · 3 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021
YearPublicationVenuePosition
2026 MFINet: Multi-view Fusion and 2D-3D Interaction Enhancement for Real-Time LiDAR Semantic Segmentation
abstract
LiDAR semantic segmentation is a key task in advanced autonomous driving systems. Projection-based methods exhibit real-time potential due to their efficiency, but suffer from inevitable 3D information loss and rely on time-consuming post-processing, limiting overall performance. To address this, we propose MFINet, a real-time semantic segmentation network based on multi-view fusion and 2D-3D interaction enhancement. It adopts a three-branch architecture that integrates 3D Point View (3D-PV), 2D Bird’s Eye View (2D-BEV) and 2D Range View (2D-RV) to make full use of 2D and 3D representation. From 3D to 2D, we design a 3D Point Feature Projector (3DPFP), which injects 3D features into the 2D BEV and RV pseudo-images to retain effective 3D information. From 2D to 3D, a Feature Enhancement (FE) module is designed to leverage the advantages of 2D information in extracting geometric and semantic features. We also introduce a 2D-3D Fusion Head (FH) to aggregate point features from multiple views. Besides, we incorporate a Multi-Scale Dilated Attention (MSDA) module with a sliding window strategy to enhance feature discrimination. Extensive experiments on the SemanticKITTI and NuScenes benchmarks demonstrate that MFINet outperforms existing methods on the SemanticKITTI, NuScenes val set and achieves competitive results on the NuScenes test set.
Nan Ma 0012, Zhijie Liu 0002, Yiheng Han
AAAI3
2026 THTFormer: Topology-adaptive hypergraph transformer network for skeleton-based action recognition
Nan Ma 0012, Genbao Xu, Yiheng Han, Beining Sun
Pattern Recognit.3
2025 Kinematic Enhanced Hypergraph Convolutional Network for Skeleton-based Human Action Recognition with LLM Training Guides
abstract
Skeleton-based human action recognition has wide applications in video understanding and virtual reality. However, most existing methods focus excessively on spatial location and global movement, while underrepresenting subtle and local actions. To address the limitation, we innovatively propose a Kinematic Enhanced Hypergraph Convolutional Network(KEHCN) with LLM training guides. The network mainly consists of LLM Training Guides(LTG), Kinematic Hypergraph Convolution(KHC), and Kinematic Gating Module(KGM). Specifically, we use the hypergraph convolutional network to extract high-order correlated human skeleton features, the KHC to encode the kinematic features and the LTG to provide a pre-trained large language model to generate text and kinematic description features during the training phase. Based on the Mixture of Experts (MoE) framework, we simplify the gating network by introducing a kinematic feature threshold, thereby constructing a dual-branch global and local motion expert network (KGM). We integrated kinematic features into KHC, LTG and KGM to seek improvements from three perspectives, all of which have enhanced the performance. The experiments on three benchmark datasets(NTU RGB+D, NTU-RGB+D 120 and NW-UCLA), demonstrate the state-of-the-art performance compared to current open-source methods.
Nan Ma 0008, Beining Sun, Yiheng Han, Genbao Xu
ACM Multimedia3
2025 KDP-MHL: Key data point-aware multi-scale hypergraph learning framework for multivariate time series classification
Nan Ma 0012, Jiacheng Guo, Yajue Yang, Shuling Li, Yiheng Han
Knowl. Based Syst.6
2025 Autonomous Tomato Harvesting With Top-Down Fusion Network for Limited Data
abstract
Using robots for tomato truss harvesting represents a promising approach to agricultural production. However, incomplete acquisition of perception information and clumsy operations often result in low harvest success rates or crop damage. To address this issue, we designed a new method for tomato truss perception, an autonomous harvesting method, and a novel circular rotary cutting end-effector. The robot performs object detection and keypoint detection on tomato trusses using the proposed Top-down Fusion Network, making decisions on suitable targets for harvesting based on phenotyping and pose estimation. The designed end-effector moves gradually from the bottom up to wrap around the tomato truss, cutting the peduncle to complete the harvest. Experiments conducted in real-world scenarios for robotic perception and autonomous harvesting of tomato trusses show that the proposed method increases accuracy by up to 11.42% and 22.29% for complete and limited dataset conditions, compared to baseline models. Furthermore, we have implemented an automatic tomato harvesting system based on TDFNet, which reaches an average harvest success rate of 89.58% in the greenhouse.
Xingxu Li, Yiheng Han, Nan Ma 0012, Yong-Jin Liu 0001, Jia Pan 0001, Siyi Zheng
IEEE Trans. Robotics2
2025 PCKRF: Point Cloud Completion and Keypoint Refinement With Fusion Data for 6D Pose Estimation
abstract
Some robust point cloud registration approaches with controllable pose refinement magnitude, such as ICP and its variants, are commonly used to improve 6D pose estimation accuracy. However, the effectiveness of these methods gradually diminishes with the advancement of deep learning techniques and the enhancement of initial pose accuracy, primarily due to their lack of specific design for pose refinement. In this paper, we propose Point Cloud Completion and Keypoint Refinement with Fusion Data (PCKRF), a new pose refinement pipeline for 6D pose estimation. The pipeline consists of two steps. First, it completes the input point clouds via a novel pose-sensitive point completion network. The network uses both local and global features with pose information during point completion. Then, it registers the completed object point cloud with the corresponding target point cloud by our proposed Color supported Iterative KeyPoint (CIKP) method. The CIKP method introduces color information into registration and registers a point cloud around each keypoint to increase stability. The PCKRF pipeline can be integrated with existing popular 6D pose estimation methods, such as the full flow bidirectional fusion network, to further improve their pose estimation accuracy. Experiments demonstrate that our method exhibits superior stability compared to existing approaches when optimizing initial poses with relatively high precision. Notably, the results indicate that our method effectively complements most existing pose estimation techniques, leading to improved performance in most cases. Furthermore, our method achieves promising results even in challenging scenarios involving textureless and symmetrical objects.
Yiheng Han, Irvin Haozhe Zhan, Long Zeng 0001, Yu-Ping Wang 0001, Ran Yi 0002, Minjing Yu, Matthieu Lin, Jenny Sheng, Yong-Jin Liu 0001
IEEE Trans. Vis. Comput. Graph.1
2024 AHPPEBot: Autonomous Robot for Tomato Harvesting based on Phenotyping and Pose Estimation
abstract
To address the limitations inherent to conventional automated harvesting robots specifically their suboptimal success rates and risk of crop damage, we design a novel bot named AHPPEBot which is capable of autonomous harvesting based on crop phenotyping and pose estimation. Specifically, In phenotyping, the detection, association, and maturity estimation of tomato trusses and individual fruits are accomplished through a multi-task YOLOv5 model coupled with a detectionbased adaptive DBScan clustering algorithm. In pose estimation, we employ a deep learning model to predict seven semantic keypoints on the pedicel. These keypoints assist in the robot’s path planning, minimize target contact, and facilitate the use of our specialized end effector for harvesting. In autonomous tomato harvesting experiments conducted in commercial green-houses, our proposed robot achieved a harvesting success rate of 86.67%, with an average successful harvest time of 32.46 s, showcasing its continuous and robust harvesting capabilities. The result underscores the potential of harvesting robots to bridge the labor gap in agriculture.
Xingxu Li, Nan Ma 0012, Yiheng Han, Siyi Zheng
ICRA3
2024 FF-LOGO: Cross-Modality Point Cloud Registration with Feature Filtering and Local to Global Optimization
abstract
Cross-modality point cloud registration is confronted with significant challenges due to inherent differences in modalities between sensors. To deal with this problem, we propose FF-LOGO: a cross-modality point cloud registration framework with Feature Filtering and LOcal-Global Optimization. The cross-modality feature correlation filtering module extracts geometric transformation-invariant features from cross-modality point clouds and achieves point selection by feature matching. We also introduce a cross-modality optimization process, including a local adaptive key region aggregation module and a global modality consistency fusion optimization module. Experimental results demonstrate that our two-stage optimization significantly improves the registration accuracy of the feature association and selection module. Our method achieves a substantial increase in recall rate compared to the current state-of-the-art methods on the 3DCSR dataset, improving from 40.59% to 75.74%. Our code will be available at https://github.com/wangmohan17/FFLOGO.
Nan Ma 0012, Yiheng Han, Yong-Jin Liu 0001
ICRA3
2024 PVP-Recon: Progressive View Planning via Warping Consistency for Sparse-View Surface Reconstruction
abstract
Neural implicit representations have revolutionized dense multi-view surface reconstruction, yet their performance significantly diminishes with sparse input views. A few pioneering works have sought to tackle this challenge by leveraging additional geometric priors or multi-scene generalizability. However, they are still hindered by the imperfect choice of input views, using images under empirically determined viewpoints. We propose PVP-Recon , a novel and effective sparse-view surface reconstruction method that progressively plans the next best views to form an optimal set of sparse viewpoints for image capturing. PVP-Recon starts initial surface reconstruction with as few as 3 views and progressively adds new views which are determined based on a novel warping score that reflects the information gain of each newly added view. This progressive view planning progress is interleaved with a neural SDF-based reconstruction module that utilizes multi-resolution hash features, enhanced by a progressive training scheme and a directional Hessian loss. Quantitative and qualitative experiments on three benchmark datasets show that our system achieves high-quality reconstruction with a constrained input budget and outperforms existing baselines.
Matthieu Lin, Jenny Sheng, Ruoyu Fan, Yiheng Han, Yubin Hu 0001, Ran Yi 0002, Yu-Hui Wen, Yong-Jin Liu 0001, Wenping Wang 0001
ACM Trans. Graph.6
2022 A Double Branch Next-Best-View Network and Novel Robot System for Active Object Reconstruction
abstract
Next best view (NBV) is a technology that finds the best view sequence for sensor to perform scanning based on partial information, which is the core part for robot active reconstruction. Traditional works are mostly based on the evaluation of candidate views through time-consuming volu-metric transformation and ray casting, which heavily limits the applications of NBV. Recent deep learning based NBV methods aim to approximately learn the evaluation function by large-scale training, and improve both the effectiveness and efficiency of NBV. However, these methods force the network to regress the exact groundtruth value of each candidate view, which is much harder than simply ranking all the candidate views. Besides, most previous NBV works assume perfect sensing and perform in simulation environments, lacking real application abilities. In this paper, we propose a novel double branch NBV network, DB-NBV, to utilize the ranking process together with the evaluation process. We further design a real NBV robot and a pipeline to conduct real active reconstruction. Experiments on both simulation and real robot show that our method achieves the best performance and can be applied to real application with high accuracy and speed.
Yiheng Han, Irvin Haozhe Zhan, Wang Zhao 0001, Yong-Jin Liu 0001
ICRA1
2021 Efficient SE(3) Reachability Map Generation via Interplanar Integration of Intra-planar Convolutions
abstract
Convolution has been used for fast computation of reachability maps, but it has high computational costs when performing SE(3) convolution operations for general joint arrangements in industrial robots and 3D workspace. Its application is also limited to planar robots, 2D workspace, or robots with special spatial arrangements for joints. In this paper, we find that the SE(3) convolution can be decomposed into a set of SE(2) convolutions, which significantly reduces the computational complexity when computing the reachability map of high-DOF robotic manipulators in the 3D workspace. We also leverage GPU parallel computing and Fast Fourier transform to further accelerate the computation procedure. We demonstrate the time efficiency and quality of our approach using a set of numerical experiments for constructing reachability maps and also present a multi-robot plant phenotyping system that uses the computed reachability map for efficient viewpoint selection and path planning.
Yiheng Han, Jia Pan 0001, Mengfei Xia, Long Zeng 0001, Yong-Jin Liu 0001
ICRA1
2020 Configuration Space Decomposition for Learning-based Collision Checking in High-DOF Robots
abstract
Motion planning for robots of high degrees-of-freedom (DOFs) is an important problem in robotics with sampling-based methods in configuration space $\mathcal{C}$ as one popular solution. Recently, machine learning methods have been introduced into sampling-based motion planning methods, which train a classifier to distinguish collision free subspace from in-collision subspace in $\mathcal{C}$. In this paper, we propose a novel configuration space decomposition method and show two nice properties resulted from this decomposition. Using these two properties, we build a composite classifier that works compatibly with previous machine learning methods by using them as the elementary classifiers. Experimental results are presented, showing that our composite classifier outperforms state-of-the-art single-classifier methods by a large margin. A real application of motion planning in a multi-robot system in plant phenotyping using three UR5 robotic arms is also presented.
Yiheng Han, Wang Zhao 0001, Jia Pan 0001, Yong-Jin Liu 0001
IROS1
2020 Ranking-Preserving Cross-Source Learning for Image Retargeting Quality Assessment
abstract
Image retargeting techniques adjust images into different sizes and have attracted much attention recently. Objective quality assessment (OQA) of image retargeting results is often desired to automatically select the best results. Existing OQA methods train a model using some benchmarks (e.g., RetargetMe), in which subjective scores evaluated by users are provided. Observing that it is challenging even for human subjects to give consistent scores for retargeting results of different source images (diff-source-results), in this paper we propose a learning-based OQA method that trains a General Regression Neural Network (GRNN) model based on relative scores-which preserve the ranking-of retargeting results of the same source image (same-source-results). In particular, we develop a novel training scheme with provable convergence that learns a common base scalar for same-source-results. With this source specific offset, our computed scores not only preserve the ranking of subjective scores for same-source-results, but also provide a reference to compare the diff-source-results. We train and evaluate our GRNN model using human preference data collected in RetargetMe. We further introduce a subjective benchmark to evaluate the generalizability of different OQA methods. Experimental results demonstrate that our method outperforms ten representative OQA methods in ranking prediction and has better generalizability to different datasets.
Yong-Jin Liu 0001, Yiheng Han, Zipeng Ye, Yukun Lai
IEEE Trans. Pattern Anal. Mach. Intell.2
2018 CFD: A Collaborative Feature Difference Method for Spontaneous Micro-Expression Spotting
abstract
Micro-expression (ME) is a special type of human expression which can reveal the real emotion that people want to conceal. Spontaneous ME (SME) spotting is to identify the subsequences containing SMEs from a long facial video. The study of SME spotting has a significant importance, but is also very challenging due to the fact that in real-world scenarios, SMEs may occur along with normal facial expressions and other prominent motions such as head movements. In this paper, we improve a state-of-the-art SME spotting method called feature difference analysis (FD) in the following two aspects. First, FD relies on a partitioning of facial area into uniform regions of interest (ROIs) and computing features of a selected sequence. We propose a novel evaluation method by utilizing the Fisher linear discriminant to assign a weight for each ROI, leading to more semantically meaningful ROIs. Second, FD only considers two features (LBP and HOOF) independently. We introduce a state-of-the-art MDMO feature into FD and propose a simple yet efficient collaborative strategy to work with two complementary features, i.e., LBP characterizing texture information and MDMO characterizing motion information. We call our improved FD method collaborative feature difference (CFD). Experimental results on two well-established SME datasets SMIC-E and CASME II show that CFD significantly improves the performance of the original FD.
Yiheng Han, Bing-Jun Li, Yukun Lai, Yong-Jin Liu 0001
ICIP1
2017 An adaptive scheduling algorithm for heterogeneous Hadoop systems
abstract
The MapReduce framework and its open source implementation Hadoop have established themselves as one of the most polular large data sets analyzers. They are widely used by many cloud service providers such as Amazon EC2 Cloud. However, while latency-sensitive applications becoming more and more important, Hadoop system shows its shortcoming in ensuring jobs completed on time. And currently, user has to provide a metric to evaluate the performance of different clients. Motivated by this, we proposed an algorithm CP-Scheduler (CPS) which uses a optimizer to analyze the best schedule in order to minimize the number of delayed jobs. Otherwise, as Hadoop System is not good at heterogeneous computing, our algorithm can also adapt different remote machines. These two features make it having better efficiency than the scheduler in Hadoop. The proposed algorithm is initially evaluated by a simulator which is designed for Hadoop. Experimental results show that the number of missing deadline jobs decrease by 60 percent on average in different sizes of situations.
Jiazhen Han, Zhengheng Yuan, Yiheng Han, Jing Liu 0012, Guangli Li
ICIS3