Haiyue Zhu

dblp:187/3714 · DBLP profile ↗
← Back
30ranked-venue papers
5as first author
25since 2021 · last 2026
0000-0002-3177-8195ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 17 · 3 first-author · 13 since 2021Artificial intelligence and machine learning · 16 · 1 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021Human-computer interaction and ubiquitous computing · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author · 2 since 2021
YearPublicationVenuePosition
2026 EPSegFZ: Efficient Point Cloud Semantic Segmentation for Few- and Zero-Shot Scenarios with Language Guidance
abstract
Recent approaches for few-shot 3D point cloud semantic segmentation typically require a two-stage learning process, i.e., a pre-training stage followed by a few-shot training stage. While effective, these methods face overreliance on pre-training, which hinders model flexibility and adaptability. Some models tried to avoid pre-training yet failed to capture ample information. In addition, current approaches focus on visual information in the support set and neglect or do not fully exploit other useful data, such as textual annotations. This inadequate utilization of support information impairs the performance of the model and restricts its zero-shot ability. To address these limitations, we present a novel pre-training-free network, named Efficient Point Cloud Semantic Segmentation for Few- and Zero-shot scenarios. Our EPSegFZ incorporates three key components. A Prototype-Enhanced Registers Attention (ProERA) module and a Dual Relative Positional Encoding (DRPE)-based cross-attention mechanism for improved feature extraction and accurate query-prototype correspondence construction without pre-training. A Language-Guided Prototype Embedding (LGPE) module that effectively leverages textual information from the support set to improve few-shot performance and enable zero-shot inference.Extensive experiments show that our method outperforms the state-of-the-art method by 5.68% and 3.82% on the S3DIS and ScanNet benchmarks, respectively.
Haiyue Zhu, Haoren Guo, Abdullah Al Mamun 0002, Cheng Xiang 0001, Tong Heng Lee
AAAI2
2025 Enhancing Multivariate Time-Series Domain Adaptation via Contrastive Frequency Graph Discovery and Language-Guided Adversary Alignment
abstract
Unsupervised domain adaptation (UDA) is a machine learning approach designed to minimize reliance on labeled data by aligning features between a labeled source domain and an unlabeled target domain, thereby reducing feature discrepancies, which is efficient for multivariate time series (MTS) prediction. However, most MTS UDA methods focus solely on aligning intra-series temporal features, overlooking the valuable information in inter-series dependencies. Research has highlighted that analyzing decomposed frequency dependencies in time series can reveal significant trends, noise patterns, and intricate temporal details. To address these unexplored frequency dependencies, we introduce the Frequency Graph Discovery Module (FGD), which uncovers and aligns shared frequency information and correlations across domains. Additionally, we propose a Frequency-Contextual Contrastive Learning (FCCL) framework to better capture and align frequency-contextual representations in multivariate time series, ensuring the extraction of label-invariant information for prediction. Furthermore, considering existing models overlooking the valuable and abundant information outside source and target dataset, we enhance the MTS UDA prediction model with a Language-guided Adversary Alignment (LAA) module, which leverages the advancement and capabilities of Large Language Models (LLMs) to get text-encoded labeled embeddings and align the classification features, thereby improving prediction accuracy. Our model achieves state-of-the-art results on three public multivariate time-series datasets for unsupervised domain adaptation, as demonstrated by empirical evidence.
Haoren Guo, Haiyue Zhu, Prahlad Vadakkepat, Weng Khuen Ho, Tong Heng Lee
AAAI2
2025 HFD-Teacher: High-Frequency Depth Distillation From Depth Foundation Models for Enhanced Depth Completion
Anqi Cheng, Haiyue Zhu, Pey Yuen Tao, Kezhi Mao
ICCV3
2025 Nonlinear Control for Underactuated Overhead Crane Using Composite Outputs under Constraints
abstract
This study presents a nonlinear feedback regulator for three-dimensional overhead cranes that exploits composite outputs to deliver potent sway suppression. Dedicated barrier functions confine these composite signals within set limits. Owing to the simple structure, the control scheme maintains oscillation suppression under velocity and cable length uncertainties. We validate stability through a Lyapunov-based proof augmented by LaSalle’s invariance principle. Simulation results confirm that the controller achieves trolley positioning and effectively cancels oscillations under external disturbances.
Shengzeng Zhang, Xinggao Liu, Michael V. Basin, Haiyue Zhu, Chentao Han, Xiongxiong He
IECON4
2025 Adaptive anti-sway control for 3D overhead crane with constraints on trolley motion and payload sway
abstract
This study proposes a nonlinear regulation controller for 3D overhead cranes, capable of achieving payload sway suppression under complex conditions. Utilizing a special form of barrier functions, constraints on the trolley motion and payload sway are derived. By analyzing the nonlinear terms, parameter estimation is incorporated to the control law, eliminating the need for prior knowledge of crane dynamics. To ensure smooth operation under varying transportation distances, a saturation function is employed to constrain the control torque generated by regulation errors. Within the Lyapunov framework, LaSalle’s invariance principle is invoked to demonstrate asymptotic convergence of the system states. Simulations validate the theoretical claims, including robustness under different transfer scenarios.
Shengzeng Zhang, Xinggao Liu, Michael V. Basin, Haiyue Zhu, Xiongxiong He
IECON4
2025 Task-Guided and Object-Centric Conditioning for Effective and Adaptive Diffusion Policy
abstract
Imitation learning has emerged as an effective paradigm for training visuo-motor policies in robotic manipulation. In real-world scenarios, visuo-motor policies are required to be effective, sample-efficient, and capable of adapting to dynamic environments. A key factor influencing these capabilities is the quality of visual representations. Conventional approaches that learn a vision encoder and policy network from scratch often result in suboptimal representations, as the training process tends to prioritize policy optimization over rich semantic feature extraction. Alternatively, while pre-trained large vision models offer strong general-purpose features, they often fail to capture the fine-grained, task-specific information required for effective manipulation. To capture rich and informative visual features, we propose TOC-DP, a novel framework that integrates SlotAttention to facilitate object-centric representation learning. Task-specific segmentation priors are incorporated as an inductive bias to enhance the task-awareness and object-awareness of the learned visual features. The extracted representations are subsequently refined to encode action-aware information during downstream policy learning. Extensive experiments on the Meta-World benchmark and real-world tasks demonstrate that TOC-DP achieves a 30% improvement in success rate over baseline methods during deployment for a variety of scenarios.
Wenshuo Wang 0004, Ruiteng Zhao, Tat Joo Teo, Marcelo H. Ang, Haiyue Zhu
IROS5
2025 SingRef6D: Monocular Novel Object Pose Estimation with a Single RGB Reference
abstract
Recent 6D pose estimation methods demonstrate notable performance but still face some practical limitations. For instance, many of them rely heavily on sensor depth, which may fail with challenging surface conditions, such as transparent or highly reflective materials. In the meantime, RGB-based solutions provide less robust matching performance in low-light and texture-less scenes due to the lack of geometry information. Motivated by these, we propose **SingRef6D**, a lightweight pipeline requiring only a **single RGB** image as a reference, eliminating the need for costly depth sensors, multi-view image acquisition, or training view synthesis models and neural fields. This enables SingRef6D to remain robust and capable even under resource-limited settings where depth or dense templates are unavailable. Our framework incorporates two key innovations. First, we propose a token-scaler-based fine-tuning mechanism with a novel optimization loss on top of Depth-Anything v2 to enhance its ability to predict accurate depth, even for challenging surfaces. Our results show a 14.41% improvement (in $\delta_{1.05}$) on REAL275 depth prediction compared to Depth-Anything v2 (with fine-tuned head). Second, benefiting from depth availability, we introduce a depth-aware matching process that effectively integrates spatial relationships within LoFTR, enabling our system to handle matching for challenging materials and lighting conditions. Evaluations of pose estimation on the REAL275, ClearPose, and Toyota-Light datasets show that our approach surpasses state-of-the-art methods, achieving a 6.1% improvement in average recall.
Haiyue Zhu, Haoren Guo, Abdullah Al Mamun 0002, Cheng Xiang 0001, Tong Heng Lee
NeurIPS2
2025 SDSimPoint: Shallow-Deep Similarity Learning for Few-Shot Point Cloud Semantic Segmentation
abstract
Three-dimensional point cloud semantic segmentation is a fundamental task in computer vision. As the fully supervised approaches suffer from the generalization issue with limited data, few-shot point cloud segmentation models have been proposed to address the flexible adaptation. Nevertheless, due to the class-agnostic nature of the few-shot pretraining, its pretrained feature extractor is hard to capture the class-related intrinsic and abstract information. Therefore, we introduce the new concept of shallow and deep similarities and propose a shallow-deep similarity learning network (SDSimPoint) that aims to learn both shallow (superficial geometry, color, etc.) and deep similarities (intrinsic context and semantics, etc.) between the support and query samples, thereby boosting the performance. Moreover, we design a beyond-episode attention module (BEAM) to enlarge the region of the attention mechanism from a single episode to the entire dataset by utilizing the memory units, which enhances the extraction ability to better capture the shallow and deep similarities. Furthermore, our distance metric function is learnable in the proposed framework, which can better adapt to complex data distributions. Our proposed SDSimPoint consistently demonstrates substantial improvements compared to baseline approaches across various datasets in diverse few-shot point cloud semantic segmentation settings.
Haiyue Zhu, Haoren Guo, Abdullah Al Mamun 0002, Cheng Xiang 0001, Clarence W. de Silva, Tong Heng Lee
IEEE Trans. Neural Networks Learn. Syst.2
2025 Low-Shot Unsupervised Visual Anomaly Detection via Sparse Feature Representation
abstract
Visual anomaly detection is an essential component in modern industrial manufacturing. Existing studies using notions of pairwise similarity distance between a test feature and nominal features have achieved great breakthroughs. However, the absolute similarity distance lacks certain generalizations, making it challenging to extend the comparison beyond the available samples. This limitation could potentially hamper anomaly detection performance in scenarios with limited samples. This article presents a novel sparse feature representation anomaly detection (SFRAD) framework, which formulates the anomaly detection as a sparse feature representation problem; and notably proposes an anomaly score by orthogonal matching pursuit (ASOMP) as a novel detection metric. Specifically, SFRAD calculates the Gaussian kernel distance between the test feature and its sparse representation in the nominal feature space for anomaly detection. Here, the orthogonal matching pursuit (OMP) algorithm is adopted to achieve the sparse feature representation. Moreover, to construct a low-redundancy memory bank storing the basis features for sparse representation, a novel basis feature sampling (BFS) algorithm is proposed by considering both the maximum coverage and the optimum feature representation simultaneously. As a result, SFRAD incorporates both the advantages of absolute similarity and linear representation; and this enhances the generalization in low-shot scenarios. Extensive experiments on the MVTec anomaly detection (MVTec AD), Kolektor surface-defect dataset (KolektorSDD), Kolektor surface-defect dataset 2 (KolektorSDD2), MVTec logical constraints anomaly detection (MVTec LOCO AD), Visual anomaly (VISA), Modified national institute of standards and technology (MNIST), and CIFAR-10 datasets demonstrate that our proposed SFRAD outperforms the previous methods and achieves state-of-the-art unsupervised anomaly detection performance. Notably, significantly improved outcomes and results have also been achieved on low-shot anomaly detection. Code is available at https://github.com/fanghuisky/SFRAD.
Fanghui Zhang, Haiyue Zhu, Yi-Gang Cen, Shichao Kan, Linna Zhang, Prahlad Vadakkepat, Tong Heng Lee
IEEE Trans. Neural Networks Learn. Syst.2
2024 GAM-Depth: Self-Supervised Indoor Depth Estimation Leveraging a Gradient-Aware Mask and Semantic Constraints
abstract
Self-supervised depth estimation has evolved into an image reconstruction task that minimizes a photometric loss. While recent methods have made strides in indoor depth estimation, they often produce inconsistent depth estimation in textureless areas and unsatisfactory depth discrepancies at object boundaries. To address these issues, in this work, we propose GAM-Depth, developed upon two novel components: gradient-aware mask and semantic constraints. The gradient-aware mask enables adaptive and robust supervision for both key areas and textureless regions by allocating weights based on gradient magnitudes. The incorporation of semantic constraints for indoor self-supervised depth estimation improves depth discrepancies at object boundaries, leveraging a co-optimization network and proxy semantic labels derived from a pretrained segmentation model. Experimental studies on three indoor datasets, including NYUv2, ScanNet, and InteriorNet, show that GAM-Depth outperforms existing methods and achieves state-of-the-art performance, signifying a meaningful step forward in indoor depth estimation. Our code will be available at https://github.com/AnqiCheng1234/GAM-Depth.
Anqi Cheng, Haiyue Zhu, Kezhi Mao
ICRA3
2024 Composite Output Feedback Control of Underactuated Overhead Crane Subject to Constraints and Parameter Uncertainties
abstract
This paper proposes a nonlinear feedback control for overhead cranes that offer satisfactory performance by taking advantages of only a composite output. Particularly, the construction of a barrier function keeps the composite output between predefined boundary values, which can enhance the safety of the system. Nonetheless, the controller with simple structure ensures the stabilization of the payload despite the presence of parametric uncertainties. To substantiate the stability proof, two analytical methodologies are employed: the Lyapunov technique and LaSalle’s invariance principle. The simulation evidences efficient positioning and oscillation elimination of the controller for various uncertain parameters, large initial errors and external disturbances, without tuning the gains of each term.
Shengzeng Zhang, Xinggao Liu, Michael V. Basin, Haiyue Zhu, Xiaoxiao Mi, Xiongxiong He
IECON4
2024 RelationGrasp: Object-Oriented Prompt Learning for Simultaneously Grasp Detection and Manipulation Relationship in Open Vocabulary
abstract
Autonomous robotic grasping under complex, clustered, and unstructured environments is a fundamental but challenging task. To achieve human-like rationality in dealing with the grasping task, the agent requires hybrid intelligence from multilateral aspects. This paper introduces RelationGrasp, a unified framework employing a transformer encoder-decoder structure to simultaneously achieve open-vocabulary object detection, manipulation relationship inference, and grasp pose detection. A unique object-oriented prompt learning mechanism is designed to seamlessly bridge the grasp pose and manipulation relationship branches, delivering high fidelity of object-grasp affiliation for object-aware grasping and grasp sequence planning. By formulating the relationship detection as an adjacency matrix regression task under multi-task learning, our framework significantly increases the relationship accuracy with reduced computational overhead. Moreover, to facilitate the robust and adaptive deployment of the proposed RelationGrasp to novel environments, we propose a consistency-based self-supervised adaptation strategy to adapt the pre-trained network to new scenarios and improve grasp accuracy on unseen objects. Our proposed network achieved state-of-the-art performance on various public dataset such as VMRD, OCID, etc., in both grasp detection and manipulation relationship classification, and real-world robot experiments has also been conducted to show the practical usages.
Songting Liu, Tat Joo Teo, Haiyue Zhu
IROS4
2024 GraspContrast: Self-supervised Contrastive Learning with False Negative Elimination for 6-DoF Grasp Detection
abstract
Robotic manipulation is a grand domain that primarily involves the use of robotic arms to interact with objects in the environment. While proposed methods have achieved advancements in grasping objects, they rely heavily on extensive training data that presents a significant challenge due to the labor-intensive process of human annotation. To address the issue, we propose GraspContrast, a self-supervised contrastive learning framework leveraging unlabeled RGB-D images to enhance point-wise feature representations for 6-DoF grasp detection. Our method designs a dual-branch network architecture to learn transformations that embed positive point pairs nearby, while pushing negative point pairs far apart. Specifically, we discuss a false negative elimination strategy to explicitly detect and remove the false negative samples that undesirably repel the point instances from the geometrically similar samples. Our method exhibits consistent improvements over existing learning-based grasp detection methods on both the GraspNet-1B benchmark and physical UR10e platform. These significant performance gains demonstrate the effectiveness of our proposed framework.
Wenshuo Wang 0004, Haiyue Zhu, Marcelo H. Ang
IROS2
2023 Few-Shot Point Cloud Semantic Segmentation via Contrastive Self-Supervision and Multi-Resolution Attention
abstract
This paper presents an effective few-shot point cloud semantic segmentation approach for real-world applications. Existing few-shot segmentation methods on point cloud heavily rely on the fully-supervised pretrain with large annotated datasets, which causes the learned feature extraction bias to those pretrained classes. However, as the purpose of few-shot learning is to handle unknown/unseen classes, such class-specific feature extraction in pretrain is not ideal to generalize into new classes for few-shot learning. Moreover, point cloud datasets hardly have a large number of classes due to the annotation difficulty. To address these issues, we propose a contrastive self-supervision framework for few-shot learning pretrain, which aims to eliminate the feature extraction bias through class-agnostic contrastive supervision. Specifically, we implement a novel contrastive learning approach with a learnable augmentor for a 3D point cloud to achieve point-wise differentiation, so that to enhance the pretrain with managed overfitting through the self-supervision. Furthermore, we develop a multi-resolution attention module using both the nearest and farthest points to extract the local and global point information more effectively, and a center-concentrated multi-prototype is adopted to mitigate the intra-class sparsity. Comprehensive experiments are conducted to evaluate the proposed approach, which shows our approach achieves state-of-the-art performance. Moreover, a case study on practical CAM/CAD segmentation is presented to demonstrate the effectiveness of our approach for real-world applications.
Haiyue Zhu, Haoren Guo, Abdullah Al Mamun 0002, Cheng-Xiang Wang 0001, Tong Heng Lee
ICRA2
2023 Lightweight Compressed Temporal and Compressed Spatial Attention with Augmentation Fusion in Remaining Useful Life Prediction
abstract
Data-driven models for predicting the Remaining Useful Lifetime (RUL) have gained popularity due to their efficiency to enhance industrial security and reduce economic losses. Recently, there has been a notable rise in the research of transformer-based models for RUL prediction. While transformer-based models have shown significant improvements over previous LSTM-based and CNN-based models, we have raised concerns regarding high computational complexity, in-effective training with low data, no sensitivity to the order of the time series, and permutation invariant on its application to RUL prediction. The persistent issue of data scarcity and the importance of capturing the temporal relations in RUL prediction further question the suitability of transformer-based models. Considering these, We propose a simple non-transformer model, Compressed Temporal and Compressed Spatial (CTCS) Attention, which is efficient and lightweight, to capture both temporal and spatial information with the incorporation of pre- and post-positional encodings. Additionally, we introduce an Augmentation Fusion Module (AFM) to enhance the comprehension ability of the invariant characteristics of the data. The proposed methodology is evaluated on the NASA Commercial Modular Aero-Propulsion System Simulation (C-MAPSS) dataset and comprehensive experiments show that our proposed method not only surpasses other methods but outperforms the transformer-based model while requiring significantly fewer Floating-point operations (FLOPs), up to 32 times less.
Haoren Guo, Haiyue Zhu, Prahlad Vadakkepat, Weng Khuen Ho, Clarence W. de Silva, Tong Heng Lee
IECON2
2023 Few-Shot Point Cloud Semantic Segmentation for CAM/CAD via Feature Enhancement and Efficient Dual Attention
abstract
Modern CAM/CAD workflows can benefit greatly from precise 3D semantic segmentation, which contributes to reducing the defect rate of work-pieces manufactured by computer-controlled CNCs and ultimately enhancing work efficiency. The majority of existing approaches for 3D object segmentation heavily rely on fully-supervised learning, where AI models are trained using extensive datasets with annotations. However, these models often exhibit unsatisfactory performance when confronted with scenarios characterized by high mixture but low volume. This means they struggle to accurately segment novel classes that were not encountered during training. In this paper, we introduce and formulate a noteworthy approach based on few-shot learning, which incorporates Sequential Dual Attention (SDA) and feature enhancement techniques. Our method aims to achieve effective semantic segmentation of point clouds in the context of CAM/CAD workflows. Unlike other few-shot models that solely adopt self-attention or lack attention, our SDA captures features at both the channel and spatial levels. Additionally, we design a non-parametric feature enhancement block to enhance the recognizability of each class in the feature space. Our proposed approach consistently demonstrates substantial enhancements in various few-shot point cloud semantic segmentation scenarios across two datasets, outperforming baseline methods.
Haiyue Zhu, Haoren Guo, Abdullah Al Mamun 0002, Clarence W. de Silva, Tong Heng Lee
IECON2
2023 Incremental few-shot learning via implanting and consolidating
Haiyue Zhu, Jun Ma 0008, Cheng Xiang 0001, Prahlad Vadakkepat
Neurocomputing2
2023 Robust Fixed-Order Controller Design for Uncertain Systems With Generalized Common Lyapunov Strictly Positive Realness Characterization
abstract
This article investigates the design of a robust fixed-order controller for single-input–single-output (SISO) polytopic systems with interval uncertainties, with the aim that the closed-loop stability is appropriately ensured and the performance specifications on sensitivity shaping are conformed in a specific finite frequency range. Utilizing the notion of generalized common Lyapunov strictly positive realness (CL-SPRness), the equivalence between strictly positive realness (SPRness) and strictly bounded realness (SBRness) is established; and then, the specifications on robust stability and performance are transformed into the SPRness of newly constructed systems and further characterized in the framework of linear matrix inequality (LMI) conditions. The proposed methodology avoids the tedious yet mandatory evaluations of the specifications on all vertices of the uncertain polytopic system in an explicit form. Instead, solving five LMIs exclusively suffices for ensuring the robust stability and performance regardless of the number of vertices, and thus, the typically heavy computational burden is considerably alleviated. It is also noteworthy that the proposed methodology additionally provides the necessary and sufficient conditions for this robust controller design with the consideration of a prescribed finite frequency range, and therefore, significantly less conservatism is attained in the system performance.
Jun Ma 0008, Haiyue Zhu, Xiaocong Li, Clarence W. de Silva, Tong Heng Lee
IEEE Trans. Syst. Man Cybern. Syst.2
2022 Incremental Few-Shot Object Detection for Robotics
abstract
Incremental few-shot learning is highly expected for practical robotics applications. On one hand, robot is desired to learn new tasks quickly and flexibly using only few annotated training samples; on the other hand, such new additional tasks should be learned in a continuous and incremental manner without forgetting the previous learned knowledge dramatically. In this work, we propose a novel Class-Incremental Few- Shot Object Detection (CI-FSOD) framework that enables deep object detection network to perform effective continual learning from just few-shot samples without re-accessing the previous training data. We achieve this by equipping the widely-used Faster-RCNN detector with three elegant components. Firstly, to best preserve performance on the pre-trained base classes, we propose a novel Dual-Embedding-Space (DES) architecture which decouples the representation learning of base and novel categories into different spaces. Secondly, to mitigate the catastrophic forgetting on the accumulated novel classes, we propose a Sequential Model Fusion (SMF) method, which is able to achieve long-term memory without additional storage cost. Thirdly, to promote inter-task class separation in feature space, we propose a novel regularization technique that extends the classification boundary further away from the previous classes to avoid misclassification. Overall, our framework is simple yet effective and outperforms the previous SOTA with a significant margin of 2.4 points in AP performance.
Haiyue Zhu, Sichao Tian, Jun Ma 0008, Chek Sing Teo, Cheng Xiang 0001, Prahlad Vadakkepat, Tong Heng Lee
ICRA2
2022 Masked Self-Supervision for Remaining Useful Lifetime Prediction in Machine Tools
abstract
Prediction of Remaining Useful Lifetime (RUL) in the modern manufacturing and automation workplace for machines and tools is essential in Industry 4.0. This is clearly evident as continuous tool wear, or worse, sudden machine breakdown, will lead to various manufacturing failures which would clearly cause economic loss. With the availability of deep learning approaches, the great potential and prospect of utilizing these for RUL prediction have resulted in several models which are designed (for RUL prediction) driven by operation data of manufacturing machines. Current efforts in these which are based on fully-supervised models heavily rely on the data labeled with their RULs. However, in these cases, the required RUL prediction data (i.e. the annotated and labeled data from faulty and/or degraded machines) can only be obtained after the machine break-down occurs. The scarcity of broken machines in the modern manufacturing and automation workplace in real- world situations increases the difficulty of getting such sufficient annotated and labeled data. In contrast, the data from healthy machines (and which are currently in operation) is much easier to be collected. Noting this challenge and the potential for improved effectiveness and applicability, we thus propose (and also fully develop) a method based on the idea of masked autoencoders which will utilize unlabeled data to do self-supervision. In thus the work here, a noteworthy masked self-supervised learning approach is developed and utilized; and this is designed to seek to build a deep learning model for RUL prediction by utilizing unlabeled data. The experiments to verify the effectiveness of this development are implemented on the C-MAPSS datasets (which is collected from the data from the NASA turbofan engine). The results rather clearly show that our development and approach here performs better, in both accuracy and effectiveness, for RUL prediction when compared with approaches utilizing a fully- supervised model.
Haoren Guo, Haiyue Zhu, Prahlad Vadakkepat, Weng Khuen Ho, Tong Heng Lee
INDIN2
2022 CAM/CAD Point Cloud Part Segmentation via Few-Shot Learning
abstract
3D part segmentation is an essential step in advanced CAM/CAD workflow. Precise 3D segmentation contributes to lower defective rate of work-pieces produced by the manufacturing equipment (such as computer controlled CNCs), thereby improving work efficiency and attaining the attendant economic benefits. A large class of existing works on 3D model segmentation are mostly based on fully-supervised learning, which trains the AI models with large, annotated datasets. However, the disadvantage is that the resulting models from the fully-supervised learning methodology are highly reliant on the completeness (or otherwise) of the available dataset, and its generalization ability is relatively poor to new unknown/unseen segmentation types (i.e., further additional so-called novel classes). In this work, we propose and develop a noteworthy few-shot learning-based approach for effective part segmentation in CAM/CAD; and this is designed to significantly enhance its generalization ability, and our development also aims to flexibly adapt to new segmentation tasks by using only relatively rather few samples. As a result, it not only reduces the requirements for the usually unattainable and exhaustive completeness of supervision datasets, but also improves the flexibility for real-world applications. In the development, drawing inspiration from the pertinent and interesting work described in the open literature as the attMPTI network, we propose and develop a multi-prototype approach (with self-attention mechanics) for few-shot point cloud part segmentation. As further improvement and innovation, we additionally adopt the transform net and the center loss block in the network. These characteristics serve to improve the comprehension for 3D features of the various possible instances of the whole work-piece and ensure the close distribution of the same class in feature space. Moreover, our approach stores data in the point cloud format that reduces space consumption, and which also makes the various procedures involved have significantly easier read and edit access (thus improving efficiency and effectiveness and lowering costs).
Haiyue Zhu, Haoren Guo, Abdullah Al Mamun 0002, Prahlad Vadakkepat, Tong Heng Lee
INDIN2
2022 Weight Imprinting Classification-Based Force Grasping With a Variable-Stiffness Robotic Gripper
abstract
Universal grasping for a diverse range of objects is a challenging problem in robotics, especially in the presence of mixed properties with fragile/rigid and heavy/light. Toward universal grasping, this article presents a practical and systematic grasping control framework that enables a variable stiffness gripper to handle the objects with diverse properties using a category-aware force regulation approach, termed classification-based force grasping. Under this framework, a convolutional neural network (CNN) is employed to classify the category of the grasping object, and a grasping force is determined based on the classified category through a database that records a predefined force magnitude per category. Sequentially, the gripper can be adjusted to a force-optimized stiffness, which facilitates the achievement of an accurate grasping force regulation in a large range. Technically, two novel enabling modules are developed for grasping classification and execution, respectively. First, a novel weight imprinting technique based on center-guided feature embedding is proposed for object classification. It enables the CNN to efficiently handle novel object categories using only a few samples even without retraining/fine-tuning. Second, a vision-based grasping force sensing module is developed, which takes advantage of the specifically designed variable-stiffness gripper. Its grasping force can be estimated from the deflection angle of finger flexure by the vision so that the contact force can be sensed and regulated. Remarkably, only single-source vision information is needed for both of the above modules without any additional force sensor. Experiments are conducted extensively to evaluate the performance of the proposed force grasping approach.Note to Practitioners—Robotic grasping often needs to handle novel categories of objects. As a result, frequent retraining of the classification neural network is a pain point, which is tedious and prone to overfitting with only a few samples. In this work, metric learning is introduced for grasping classification where a novel kind of weight imprinting classification is proposed to handle the novel classes by better feature embedding and directly setting the classifier weights without retraining or fine-tuning. Together with the benefits from the variable stiffness feature of the gripper, the proposed vision-based force grasping approach can handle a wide range of objects from fragile to heavy, and the grasping force is controllable from 0.2 N onward to the motor limitation. The controllable grasping force resolution of the proposed vision grasping is better than 0.05 N, the accuracy of the grasping force is evaluated from 0.2 to 12 N, and the evaluated grasping objects are from extremely fragile potato chips and eggshell to heavy flange and metal block.
Haiyue Zhu, Xiong Li 0001, Xiaocong Li, Jun Ma 0008, Chek Sing Teo, Tat Joo Teo, Wei Lin 0002
IEEE Trans Autom. Sci. Eng.1
2022 On Robust Stability and Performance With a Fixed-Order Controller Design for Uncertain Systems
abstract
Typically, it is desirable to design a control system that is not only robustly stable in the presence of parametric uncertainties but also guarantees an adequate level of system performance. However, most of the existing methods need to take all extreme models over an uncertain domain into consideration, which then results in costly computation. Also, since these approaches attempt rather unrealistically to guarantee the system performance over a full frequency range, a conservative design is always admitted. Here, taking a specific viewpoint of robust stability and performance under a stated restricted frequency range (which is applicable in rather many real-world situations), this article provides an essential basis for the design of a fixed-order controller for a system with bounded parametric uncertainties, which avoids the tedious but necessary evaluations of the specifications on all the extreme models in an explicit manner. A Hurwitz polynomial is used in the design and the robust stability is characterized by the notion of positive realness, such that the required robust stability condition is then successfully constructed. Also, the robust performance criteria in terms of sensitivity shaping under different frequency ranges are constructed based on an approach of bounded realness analysis. Furthermore, the conditions for robust stability and performance are expressed in the framework of linear matrix inequality (LMI) constraints, and thus can be efficiently solved. Comparative simulations are provided to demonstrate the effectiveness and efficiency of the proposed approach.
Jun Ma 0008, Haiyue Zhu, Masayoshi Tomizuka, Tong Heng Lee
IEEE Trans. Syst. Man Cybern. Syst.2
2021 Few-Shot Object Detection via Classification Refinement and Distractor Retreatment
abstract
We aim to tackle the challenging Few-Shot Object Detection (FSOD), where data-scarce categories are presented during the model learning. The failure modes of FasterRCNN in FSOD are investigated, and we find that the performance degradation is mainly due to the classification incapability (false positives) caused by category confusion, which motivates us to address FSOD from a novel aspect of classification refinement. Specifically, we address the intrinsic limitation from the aspects of both architectural enhancement and hard-example mining. We introduce a novel few-shot classification refinement mechanism where a decoupled Few-Shot Classification Network (FSCN) is employed to improve the final classification of a base detector. Moreover, we especially probe a commonly-overlooked but destructive issue of FSOD, i.e., the presence of distractor samples due to the incomplete annotations where images from the base set may contain novel-class objects but remain unlabelled. Retreatment solutions are developed to eliminate the incurred false positives. For FSCN training, the distractor is formulated as a semi-supervised problem, where a distractor utilization loss is proposed to make proper use of it for boosting the data-scarce classes, while a confidence-guided dataset pruning (CGDP) technique is developed to facilitate the few-shot adaptation of base detector. Experiments demonstrate that our proposed framework achieves state-of-the-art FSOD performance on public datasets, e.g., Pascal VOC and MS-COCO.
Haiyue Zhu, Chek Sing Teo, Cheng Xiang 0001, Prahlad Vadakkepat, Tong Heng Lee
CVPR2
2021 Robust Control of a Two-Degree-of-Freedom Flexure-Based Nanopositioner for Planar Scanning Tasks
abstract
A two-degree-of-freedom (2-DoF) flexure-based nanopositioner is investigated for the planar scanning tasks, and a robust controller design scheme based on the convex inner approximation method is proposed. In practice, a flexure-based mechanism is usually represented by a second-order dynamic model. However, the second-order dynamic model cannot precisely fit the real system dynamics, and the model mismatch renders it difficult to achieve satisfying system performance in applications. Such a mismatch includes the parameter uncertainties caused by inaccurate model identification, different motion conditions, as well as high-order resonances. Note that if the controller is not well designed, the high-order resonances can be frequently activated, especially when the system input variation is significant. Therefore, to deal with the above impediments, a novel scheme for the robust controller design is proposed, with the variation of system input considered. In the proposed scheme, a subset of gains that can stabilize the closed-loop system is characterized elegantly via an inner approximation method considering the model uncertainties, and the formulated optimization problem regarding the determination of the controller parameters can be efficiently solved. Furthermore, the proposed scheme guarantees the performance regarding the H2-norm level and limits the H∞-norm level in a designated range. Finally, numerical optimization and comparative experiments are carried out, and the results evidently show the effectiveness of the proposed method.
Zilong Cheng, Jun Ma 0008, Xiaocong Li, Haiyue Zhu, Tong Heng Lee
SMC7
2020 Learning-Based Controller Optimization for Repetitive Robotic Tasks
abstract
Dynamic control for robotic automation tasks is traditionally designed and optimized with a model-based approach, and the performance relies heavily upon accurate system modeling. However, modeling the true dynamics of increasingly complex robotic systems is an extremely challenging task and it often renders the automation system to operate in a non-optimal condition. Notably, many industrial robotic applications involve repetitive motions and constantly generate a large amount of motion data under the non-optimal condition. These motion data contain rich information, and therefore an intelligent automation system should be able to learn from these non-optimal motion data to drive the system to operate optimally in a data-driven manner. In this paper, we propose a learning-based controller optimization algorithm for repetitive robotic tasks. To achieve this, a multi-objective cost function is designed to take into consideration both the trajectory tracking accuracy and smoothness, and then a data-driven approach is developed to estimate the gradient and Hessian based on the motion data for optimization without relying on the dynamic model. Experiments based on a magnetically-levitated nanopositioning system are conducted to demonstrate the effectiveness and practical appeals of the proposed algorithm in repetitive robotic automation tasks.
Xiaocong Li, Haiyue Zhu, Jun Ma 0008, Tat Joo Teo, Chek Sing Teo, Masayoshi Tomizuka, Tong Heng Lee
IROS2
2020 Grasping Detection Network with Uncertainty Estimation for Confidence-Driven Semi-Supervised Domain Adaptation
abstract
Data-efficient domain adaptation with only a few labelled data is desired for many robotic applications, e.g., in grasping detection, the inference skill learned from a grasping dataset is not universal enough to directly apply on various other daily/industrial applications. This paper presents an approach enabling the easy domain adaptation through a novel grasping detection network with confidence-driven semi-supervised learning, where these two components deeply interact with each other. The proposed grasping detection network specially provides a prediction uncertainty estimation mechanism by leveraging on Feature Pyramid Network (FPN), and the mean-teacher semi-supervised learning utilizes such uncertainty information to emphasizing the consistency loss only for those unlabelled data with high confidence, which we referred it as the confidence-driven mean teacher. This approach largely prevents the student model to learn the incorrect/harmful information from the consistency loss, which speeds up the learning progress and improves the model accuracy. Our results show that the proposed network can achieve high success rate on the Cornell grasping dataset, and for domain adaptation with very limited data, the confidence- driven mean teacher outperforms the original mean teacher and direct training by more than 10% in evaluation loss especially for avoiding the overfitting and model diverging.
Haiyue Zhu, Fengjun Bai, Xiaocong Li, Jun Ma 0008, Chek Sing Teo, Pey Yuen Tao, Wei Lin 0002
IROS1
2019 Flexure-Based Magnetically Levitated Dual-Stage System for High-Bandwidth Positioning
abstract
Bandwidth is a critical specification for motion positioning systems because fast response to reference and broad-band disturbance rejection is highly desirable in industrial applications, e.g., two-dimensional (2-D)/3-D scanning. This paper presents a parallel-actuation, dual-stage concept to enhance the bandwidth of magnetically levitated (maglev) positioning system, which is realized by utilizing compliant joints to construct a monolithic-cut flexure-based fine positioning stage within the primary maglev stage, hence turning such a dual-stage system into a fully cable-less maglev system. An integrated design approach is employed to design the flexure-based secondary stage by optimizing both the mechanical parameters and the controller parameters, where various specifications, e.g., stability, performance, and saturation, are considered under the proposed framework. Experimental results have shown that the prototype can achieve a root-mean-square error of 43 nm in the principal axis even though the accuracy of the primary maglev stage is limited in micron-level because of the noise of capacitive sensors. Results also show that the developed prototype can significantly improve the closed-loop bandwidth of the maglev system from 20 to around 200 Hz.
Haiyue Zhu, Tat Joo Teo, Chee Khiang Pang
IEEE Trans. Ind. Informatics1
2015 Analysis of force harmonics and eddy current damping for 2 DOF moving magnet linear motor
abstract
This paper presents the analysis of force harmonics and eddy current damping for 2 DOF moving magnet linear motor (MMLM), which is utilized in magnetically levitated (maglev) positioning systems. A novel current-force model considering higher-order harmonics is proposed for MMLM in this paper, and this model provides an analytical tool to analyze the force ripple effect in MMLM theoretically. A commutation law is derived based on the proposed current-force model, which can ideally eliminate the force ripple in theory. In addition, the eddy current damping effect for a moving Halbach permanent magnet (PM) array on the conductive plate is analytically modeled in this work. By utilizing this model, the damping effect of MMLM can be derived using first principle. Finally, a prototype of MMLM is fabricated and experiment is conducted to implement the proposed commutation law and verify the eddy current damping model.
Haiyue Zhu, Tat Joo Teo, Chee Khiang Pang
IECON1
2014 Modeling and design of a size and mass reduced magnetically levitated planar positioner
abstract
Magnetic lévitation technology is a promising solution to achieve ultra-precision motion. This paper presents a design of maglev planar positioner with less size and mass, compared with existing designs. The maglev planar positioner employs four groups of low-order moving magnet linear motors (MMLM) to provide force, and the Halbach PM array in each MMLM contains only one magnetic pole. As a result, the size and mass of the proposed maglev planar positioner is significantly reduced. Due to the inaccuracy of existing force models in predicting such a low-order MMLM, a novel modeling approach is introduced to derive the analytical force model in this paper. Finally, a prototype of this MMLM with one magnetic pole Halbach PM array is developed, and the experiment is conducted to validate the accuracy of the proposed force modeling approach.
Haiyue Zhu, Chee Khiang Pang, Tat Joo Teo, Lubecki Tomasz Marek
IECON1