Zhi-Xin Yang 0001

dblp:22/11351-1 · also Zhixin Yang 0001 · DBLP profile ↗
← Back
70ranked-venue papers
4as first author
52since 2021 · last 2026
0000-0001-9151-7758ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 27 · 3 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 23 · 1 first-author · 19 since 2021Applied, interdisciplinary, general and emerging computing · 16 · 1 first-author · 14 since 2021Databases, data management, data science and information retrieval · 5 · 5 since 2021Computer networks · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Distributed fixed-time leader-referenced rigid shape formation control for multi-robot vehicles with prescribed performance
Zhongchao Liang, Jian Pan 0001, Yunfeng Hu 0003, Zhi-Xin Yang 0001, Jing Zhao 0010
Adv. Eng. Informatics6
2026 Robust path tracking control for four wheel independently actuated electric vehicle with probabilistic time-varying delays
Jiachen Wei, Pak-Kin Wong 0001, Zhi-Xin Yang 0001, Wenfeng Li 0002, Dawei Pi, Jing Zhao 0010
Adv. Eng. Informatics4
2026 Target-agnostic common attributes learning for few-shot semantic segmentation
Yadang Chen, Yuhui Zheng, Zhi-Xin Yang 0001, Enhua Wu
Pattern Recognit.4
2026 Adaptive Energy-Saving Control for Vehicle Suspensions via Coupling and Disturbance Effects Utilization
abstract
This paper proposes a novel energy-efficient control method for vehicle suspensions that addresses key application-oriented issues, including inevitable disturbance, inherent nonlinearity, state-coupling effects, and energy consumption. Unlike conventional schemes, the proposed method explicitly quantifies the influence of disturbance and state-coupling by designing dedicated effect indicators and exploiting the beneficial aspects of nonlinear dynamics. Positive effects are harnessed, while negative effects are transformed into advantageous contributions, thereby delivering superior robustness, lower energy consumption, and improved response rapidity. Specifically, a fuzzy disturbance observer is developed to accurately estimate the disturbance factors, including both parametric/unmolded uncertainty and external disturbance. Effect indicators are introduced to characterize the positive and negative impacts of disturbance and coupling on the active suspension system, and these insights are incorporated into the design of a novel adaptive controller. In addition, biologically inspired nonlinear reference model is deliberately introduced to make use of favorable nonlinear stiffness and damping effects. Furthermore, Lyapunov’s theory is employed to ensure the asymptotic stability of the overall suspension system. Experimental validation demonstrated that the proposed approach achieves excellent transient performance and significant energy savings, with reductions of up to 80% or more.
Menghua Zhang, Jing Zhao 0010, Haokun Geng, Zengcheng Zhou, Zhi-Xin Yang 0001
IEEE Trans Autom. Sci. Eng.5
2026 Boosting Video Object Segmentation With Discriminative Core Features and Adaptive Position Refinement
Yadang Chen, Guolong Li, Yuhui Zheng, Bin Sheng 0001, Zhi-Xin Yang 0001, Enhua Wu
IEEE Trans. Circuits Syst. Video Technol.5
2026 TiGDistill-BEV: Multi-View BEV 3D Object Detection via Target Inner-Geometry Learning Distillation
abstract
Accurate multi-view 3D object detection is essential for applications such as autonomous driving. Researchers have consistently aimed to leverage LiDAR’s precise spatial information to enhance camera-based detectors through methods like depth supervision and bird-eye-view (BEV) feature distillation. However, existing approaches often face challenges due to the inherent differences between LiDAR and camera data representations. In this paper, we introduce the TiGDistill-BEV, a novel approach that effectively bridges this gap by leveraging the strengths of both sensors. Our method distills knowledge from diverse modalities(e.g., LiDAR) as the teacher model to a camera-based student detector, utilizing the Target Inner-Geometry learning scheme to enhance camera-based BEV detectors through both depth and BEV features by leveraging diverse modalities. Specially, we propose two key modules: an inner-depth supervision module to learn the low-level relative depth relations within objects which equips detectors with a deeper understanding of object-level spatial structures, and an inner-feature BEV distillation module to transfer high-level semantics of different keypoints within foreground targets. To further alleviate the domain gap, we incorporate both inter-channel and inter-keypoint distillation to model feature similarity. Extensive experiments on the nuScenes benchmark demonstrate that TiGDistill-BEV significantly boosts camera-based only detectors achieving a state-of-the-art with 62.8% NDS and surpassing previous methods by a significant margin. The codes is available at: https://github.com/Public-BOTs/TiGDistill-BEV.git.
Shaoqing Xu, Peixiang Huang, Ziying Song, Zhi-Xin Yang 0001
IEEE Trans. Circuits Syst. Video Technol.5
2026 Double Diffusion Policy for Robust Robot Learning via Human Guidance
abstract
The diffusion policy has introduced the excellent multimodal data modeling capability of the diffusion model to robotic imitation learning. However, extensive expert demonstration data is still needed during training, requiring considerable human effort to collect. Natural human action data is cheaper and easier to obtain than expert demonstration data, but its use in direct robot policy training is challenging due to the distribution gap. In this study, we propose the double diffusion policy learning paradigm, which incorporates low-cost human action data to diminish the reliance of diffusion policy on extensive expert demonstration data. Specifically, we extract human intentions from the diffusion policy modeling human actions, then integrate these intentions into training the diffusion policy for generating robot actions. This processing guides the policy model to train better on small expert demonstration datasets. Experiments show that our double diffusion policy outperforms the vanilla diffusion policy and other state-of-the-art imitation learning algorithms with limited expert demonstration data.
Weixiang Liang, Ying Gong, Xiongyi Li, Yinlong Liu, Zhi-Xin Yang 0001
IEEE Trans. Ind. Informatics5
2026 Learnable Object Queries for Few-Shot Semantic Segmentation
abstract
Few-shot semantic segmentation (FSS) aims to segment unseen-category objects given only a few annotated samples. Although significant progress has been made in the field of FSS, selecting an appropriate feature matching method remains a challenge. Traditional prototype-based methods can preserve high-level semantic features, but they tend to lose detailed information. On the other hand, pixel-level comparison methods retain fine-grained details but are vulnerable to distractors and noise, leading to poor robustness. To address these issues, this paper proposes a target-agnostic object-based method. Specifically, we propose a set of learnable "object queries" to extract object features, which preserve both high-level semantic information and fine-grained details. Additionally, during the training phase, we exploit the prior knowledge of foreground and background embedded in the samples to enhance the model's performance. In the inference phase, the model utilizes both the support set and the learned prior knowledge to perform segmentation tasks, mitigating the data distribution bias caused by limited samples. Extensive experiments on benchmark datasets demonstrate that our method outperforms state-of-the-art approaches in both accuracy and robustness. Code is available at https://github.com/wenbo456/OTBNet.
Yadang Chen, Yuhui Zheng, Zhi-Xin Yang 0001, Enhua Wu
IEEE Trans. Image Process.4
2026 Refinement on Both Foreground and Background Prototype for Few-Shot Segmentation
abstract
Although few-shot segmentation (FSS) methods have achieved remarkable results, there remain challenges associated with the limited number of support samples.i)The objects in support and query images may have substantially different appearances even though they belong to the same category, which is known as the prototype bias problem.ii)Most methods neglect the background information, especially the query background during the inference stage. To address these problems, we propose DPRNet, a novel network with dual branch of foreground and background prototype refinement modules. Specifically, we first present a Variational Feature Semantic Enhancement (VFSE) module, in which we refine the object prototype with a variational autoencoder and word-text labels. In this way, the biased class-wise prototype caused by the limited support samples can be aligned, achieving better performance. Second, we design a Background Prototype Refinement (BPR) module that effectively explores the potential information in the background for both the support and query images. More importantly, it is designed to generate online predictions of the query background during the training stage to fully mimic the inference stage. These advancements enhance the robustness and generalizability of our method, and the results of experiments demonstrate its effectiveness. In the 1-way 5-shot setting on PASCAL-$5^{i}$, our method achieves a mean-IoU improvement of 1.59% over the competing method.
Yadang Chen, Yuhui Zheng, Zhi-Xin Yang 0001, Enhua Wu
IEEE Trans. Multim.4
2025 Review of intelligent fault diagnosis for rotating machinery under imperfect data conditions
Hao Chen 0099, Jiaming Li 0019, Xianbo Wang, Lianqing Yu, Zhi-Xin Yang 0001
Expert Syst. Appl.5
2025 DDConv: Disentangled Dual Branch Convolutional Network for Anomaly Detection in Server Machines
abstract
Detecting anomalies in server machines is critical for providing reliable web services. Usually, working status of server machines are characterized by multivariate time series (MTS). However, complex patterns in MTS including both temporal and inter-variable ones make it hard to precisely differentiate anomalies from normal sequences. In this paper, we propose a model called DDConv, a reconstruction-based fully convolutional network to detect anomalies in server machine MTS. The core of DDConv is to disentangle pattern learning in time and variable dimensions and separate the intra-dimension and inter-dimension feature learning processes for thoroughly exploring all these kinds of patterns. Patterns in time and variable dimensions are captured independently in parallel via a dual branch convolutional network working under the macro-group convolution mechanism. A novel regularization is further proposed to enable precise reconstructions by shrinking the error space formed by reconstruction errors. Extensive experiments on public benchmarks demonstrate DDConv achieves superior anomaly detection performance over state-of-the-art baselines while maintaining model efficiency.
Yueyue Yao, Zhi-Xin Yang 0001
IEEE Internet Things J.2
2025 Adaptive set-level metric for few-Shot image classification
Yadang Chen, Jin Wang 0005, Zhi-Xin Yang 0001
Neural Networks4
2025 Robustly solving PnL problem using Clifford tori
Yinlong Liu, Shengyong Ding, Zhi-Xin Yang 0001
Pattern Recognit.4
2025 Trajectory Progress-Based Prioritizing and Intrinsic Reward Mechanism for Robust Training of Robotic Manipulations
abstract
Training robots by model-free deep reinforcement learning (DRL) to carry out robotic manipulation tasks without sufficient successful experiences is challenging. Hindsight experience replay (HER) is introduced to enable DRL agents to learn from failure experiences. However, the HER-enabled model-free DRL still suffers from limited training performance due to its uniform sampling strategy and scarcity of reward information in the task environment. Inspired by the progress incentive mechanism in human psychology, we propose Progress Intrinsic Motivation-based HER (P-HER) in this work to overcome these difficulties. First, the Trajectory Progress-based Prioritized Experience Replay (TPPER) module is developed to prioritize sampling valuable trajectory data thereby achieving more efficient training. Second, the Progress Intrinsic Reward (PIR) module is introduced in agent training to add extra intrinsic rewards for encouraging the agents throughout the exploration of task space. Experiments in challenging robotic manipulation tasks demonstrate that our P-HER method outperforms original HER and state-of-the-art HER-based methods in training performance. Our code of P-HER and its experimental videos in both virtual and real environments are available athttps://github.com/weixiang-smart/P-HER. Note to Practitioners—This work is motivated to develop a fast and effective learning method for intelligent robotic manipulation of typical industrial tasks, including pushing, picking, and placing workpieces, which are essential and fundamental processing plan activities for accomplishing robotic machining and assembly applications towards smart manufacturing. The introduction of reinforcement learning enables robots to learn manipulation tasks autonomously, which can save the effort for engineers to teach or hard program the robot and also reduce labor costs. However, the existing HER-based reinforcement learning algorithms are with low training efficiency and performance due to the uniform sampling and scant task reward. Inspired by human learning, this work introduces a progress incentive mechanism to identify valuable trajectory data for effective training. In addition, a novel rewarding method, that applies additional intrinsic rewards for agents learning valuable trajectory space, results in fast and robust learning. The setting of important weight parameters in the rewarding method is given in the paper, which provides a practical reference for applying the proposed algorithm. The average success rate of two actual manipulation tasks in simulation and real robotic manipulation environments are 96% and 92.5%, respectively, which demonstrates that the method is effective for both environments and there is 3.5% average gap of successful rate dropping from simulation scenarios to real ones due to the inherent mismatches between simulation and reality. The high success rate demonstrated in the real Workpieces-sorting task exemplifies the potential of the trained policies for application in industrial scenarios.
Weixiang Liang, Yinlong Liu, Jikun Wang, Zhi-Xin Yang 0001
IEEE Trans Autom. Sci. Eng.4
2025 TactCLNet: Tactile Continual Learning Network Based on Generative Replay for Object Hardness Recognition
abstract
Currently, deep neural networks can be extremely effective in robotic tactile perception. However, a major challenge is to solve the problem of continual learning of robotic tactile perception in an open and dynamic environment. In this paper, we propose a novel continual learning method for the domian incremental learning task in the field of tactile perception. To be specific, we introduce a morphology-specific variational autoencoders which can mitigate catastrophic forgetting by generating pseudo-samples for training in the continual learning process. We integrate the generative model and the discriminative model into one model, which reduces the size of model and improves the continual learning ability. In addition, considering the ordinal information between the hardness levels, we propose to add conditional information to the model and introduce a modified loss function to combine the latent value with the hardness information, which improves the continual learning performance by controlling the distribution and quality of pseudo-sample generation. Following this, we designed a tactile robot experiment, collected hardness data, and tested our model on this object hardness recognition task. We show experimentally that, after training, the model can still maintain the accuracy of more than 94% after learning three tasks in terms. Note to Practitioners—In the field of robotics tactile perception, the issue of continual learning in robots is a crucial problem that urgently requires resolution. We hope robots to effectively engage in continual learning across multiple tasks, ensuring the acquisition of new knowledge while mitigating the risk of forgetting previously acquired knowledge. In this paper, we propose a novel continual learning method for the domian incremental learning task. we introduce a morphology-specific variational autoencoders based on replaying pseudo-samples during continual learning process which reduces the size of model and improves the continual learning ability. We enhance model performance by integrating generative and discriminative models, incorporating conditional information to control the distribution of replayed sample types, and leveraging sequential relationships among samples. It is proved that the proposed method is able to effectively improve the accuracy in a tactile domian incremental learning task.
Zhengkun Yi, Senlin Fang, Yupo Zhang, Feng Wan 0003, Zhi-Xin Yang 0001, Xu Lu 0002, Xinyu Wu 0001
IEEE Trans Autom. Sci. Eng.6
2025 SF-Pose: Semantic-Fusion Six Degrees of Freedom Object Pose Estimation via Pyramid Transformer for Industrial Scenarios
abstract
Object six degrees of freedom (6-DoF) pose estimation is the powerful vision algorithm for the robot-environment interaction. However, current robust pose estimation algorithms rely heavily on labeled real data with high-cost collection, making it difficult to apply the algorithm. Many studies discuss the use of synthetic data as a complement to real datasets. However, reducing the gap between synthetic and real data is still a challenging problem. Based on the consistency of object geometric characteristics between real data and synthetic data, we argue that multi-input, rather than image-only input, is more suitable for transfer from synthetic to real, because it strengthens the extraction of object geometric feature. Therefore, we propose a semantic-fusion 6-DoF object pose estimation method that effectively capture common features across various resolutions by employing the designed pyramid transformer feature-fusion module. Extensive experiments show that the proposed method performs better than the state-of-the-art (SOTA), indicating that the proposed method can effectively extract and fuse different representations. Furthermore, in response to the lack of industrial scene datasets, we also develop a synthetic pose dataset and conduct the human-robot collaboration experiment to verify the robustness of the proposed method. Note to Practitioners—The purpose of this paper is to bridge the gap between synthetic and real data for pose estimation of industrial tools. Our method can be trained only on synthetic data and accurately estimate pose parameters in real scenes. Combining physically-based renderer and industrial tools, such as hammers and screwdrivers, a synthetic dataset of industrial scenes can be produced using the data production pipeline proposed in this paper. In this case, the trained model can assist the robot vision system to understand object pose information in a real production workshop. Extensive dataset experiments and human-robot collaboration experiments demonstrate the effectiveness of the proposed method. In addition, based on the actual robot working environment, practitioners can produce industrial datasets from multiple angles, objects, and scenes. Sufficient datasets can enhance the model’s generalization and robustness.
Jikun Wang, Yinlong Liu, Zhi-Xin Yang 0001
IEEE Trans Autom. Sci. Eng.3
2025 Physical-Knowledge-Guided and Interpretable Deep Neural Networks for Gear Fault Severity Level Diagnosis
abstract
While deep-learning (DL) models have achieved significant achievements in fault diagnosis, their inherent opacity for human users often hinders practical applications in risk-sensitive scenarios. Fortunately, the advent of class activation mapping (CAM) significantly enhanced the transparency of DL models by illuminating the specific input areas that contribute more to the classification results. Nevertheless, CAM fails to enhance diagnostic accuracy and actively leverage interpretability due to its passively explanatory property for the trained models. To address this issue, in this article, a physically meaningful regularization (PMR) term is proposed by using gradient-weighted CAM, to guide the models in focusing on the same frequency bands of the input spectra and ignoring other parts of noisy and irrelevant signals. Based on the PMR term, a two-step back propagation training algorithm is accordingly designed to train the diagnostic models embedded with physical knowledge. Consequently, the obtained physical-knowledge-guided and interpretable DL models can offer not only strong interpretability but also a higher diagnostic accuracy for the noised test samples. Finally, the proposed diagnostic method is validated in two datasets containing multiple fault severity levels. The diagnostic results, along with the saliency analysis, substantiate the efficacy of the proposed method.
Jiaming Li 0019, Xianbo Wang, Hao Chen 0099, Zhi-Xin Yang 0001
IEEE Trans. Ind. Informatics4
2025 Fusion-Perception-to-Action Transformer: Enhancing Robotic Manipulation With 3-D Visual Fusion Attention and Proprioception
abstract
Most prior robot learning methods focus on image-based observations, limiting their capability in 3-D robotic manipulation. Voxel representation naturally delivers rich spatial features but remains underutilized. Specifically, current voxel-based methods struggle with fine-grained tasks, since precise actions are not fully achievable. However, humans can accomplish these tasks well using vision and proprioception. Inspired by this, this article proposed a novel Fusion-Perception-to-Action Transformer (FP2AT) with cross-layer feature aggregation to handle fine-grained manipulation in 3-D space. In particular, a multiscale 3-D visual fusion attention mechanism is devised to draw attention to local regions of interest and maintain awareness of global scenes, thereby boosting the capabilities of visual perception and action planning. Meanwhile, a 3-D visual mutual attention mechanism is designed and it can also enhance spatial perception. Besides, we further explore the potential of FP2AT by developing its coarse-to-fine version, which progressively refines the action space for more precise predictions. In addition, a proprioceptive encoder is developed to mimic the perception of body movements and contact, elevating the effectiveness of the FP2AT. Furthermore, a new metric, the average number of key actions (ANKA), is introduced to evaluate efficiency and planning capability. In various simulated and real-robot examples, our methods significantly outperform state-of-the-art 3-D-vision-based methods in success rate and ANKA metrics.
Yangjun Liu, Binghan Chen, Zhi-Xin Yang 0001, Sheng Xu 0004
IEEE Trans. Robotics4
2024 Space-time Reinforcement Network for Video Object Segmentation
abstract
Recently, video object segmentation (VOS) networks typically use memory-based methods: for each query frame, the mask is predicted by space-time matching to memory frames. Despite these methods having superior performance, they suffer from two issues: 1) Challenging data can destroy the space-time coherence between adjacent video frames. 2) Pixel-level matching will lead to undesired mismatching caused by the noises or distractors. To address the aforementioned issues, we first propose to generate an auxiliary frame between adjacent frames, serving as an implicit short-temporal reference for the query one. Next, we learn a prototype for each video object and prototype-level matching can be implemented between the query and memory. The experiment demonstrated that our network outperforms the state-of-the-art method on the DAVIS 2017, achieving a ℐ&ℱ score of 86.4%, and attains a competitive result 85.0% on YouTube VOS 2018. In addition, our network exhibits a high inference speed of 32+ FPS.
Yadang Chen, Zhi-Xin Yang 0001, Enhua Wu
ICME3
2024 A Transformer-Based Adaptive Prototype Matching Network for Few-Shot Semantic Segmentation
Yadang Chen, Yuhui Zheng, Zhi-Xin Yang 0001, Enhua Wu
IJCAI4
2024 SparseInteraction: Sparse Semantic Guidance for Radar and Camera 3D Object Detection
abstract
Multi-modal fusion techniques, such as radar and images, enable a complementary and cost-effective perception of the surrounding environment regardless of lighting and weather conditions. However, existing fusion methods for surround-view images and radar are challenged by the inherent noise and positional ambiguity of radar, which leads to significant performance losses. To address this limitation effectively, our paper presents a robust, end-to-end fusion framework dubbed SparseInteraction. First, we introduce the Noisy Radar Filter (NRF) module to extract foreground features by creatively using queried semantic features from the image to filter out noisy radar features. Furthermore, we implement the Sparse Cross-Attention Encoder (SCAE) to effectively blend foreground radar features and image features to address positional ambiguity issues at a sparse level. Ultimately, to facilitate model convergence and performance, the foreground prior queries containing position information of the foreground radar are concatenated with predefined queries and fed into the subsequent transformer-based decoder. The experimental results demonstrate that the proposed fusion strategies markedly enhance detection performance and achieve new state-of-the-art results on the nuScenes benchmark. Source code is available at https://github.com/GG-Bonds/SparseInteraction.
Shaoqing Xu, Shengyin Jiang, Li Liu 0069, Ziying Song, Zhi-Xin Yang 0001
ACM Multimedia7
2024 Wavelet Potentials: An Efficient Potential Recovery Technique for Pointwise Incompressible Fluids
abstract
Abstract We introduce an efficient technique for recovering the vector potential in wavelet space to simulate pointwise incompressible fluids. This technique ensures that fluid velocities remain divergence‐free at any point within the fluid domain and preserves local volume during the simulation. Divergence‐free wavelets are utilized to calculate the wavelet coefficients of the vector potential, resulting in a smooth vector potential with enhanced accuracy, even when the input velocities exhibit some degree of divergence. This enhanced accuracy eliminates the need for additional computational time to achieve a specific accuracy threshold, as fewer iterations are required for the pressure Poisson solver. Additionally, in 3D, since the wavelet transform is taken in‐place, only the memory for storing the vector potential is required. These two features make the method remarkably efficient for recovering vector potential for fluid simulation. Furthermore, the method can handle various boundary conditions during the wavelet transform, making it adaptable for simulating fluids with Neumann and Dirichlet boundary conditions. Our approach is highly parallelizable and features a time complexity of O(n), allowing for seamless deployment on GPUs and yielding remarkable computational efficiency. Experiments demonstrate that, taking into account the time consumed by the pressure Poisson solver, the method achieves an approximate 2x speedup on GPUs compared to state‐of‐the‐art vector potential recovery techniques while maintaining a precision level of 10−6 when single float precision is employed. The source code of ‘Wavelet Potentials’ can be found in https://github.com/yours321dog/WaveletPotentials .
Luan Lyu, Xiaohua Ren, Wei Cao 0008, Jian Zhu 0001, Enhua Wu, Zhi-Xin Yang 0001
Comput. Graph. Forum6
2024 Privacy-preserving intelligent fault diagnostics for wind turbine clusters using federated stacked capsule autoencoder
Hao Chen 0099, Xianbo Wang, Zhi-Xin Yang 0001, Jiaming Li 0019
Expert Syst. Appl.3
2024 Learning self-target knowledge for few-shot segmentation
Yadang Chen, Zhi-Xin Yang 0001, Enhua Wu
Pattern Recognit.3
2024 OA-Pose: Occlusion-aware monocular 6-DoF object pose estimation under geometry alignment for robot manipulation
Jikun Wang, Luqing Luo, Weixiang Liang, Zhi-Xin Yang 0001
Pattern Recognit.4
2024 Boosting Video Object Segmentation via Robust and Efficient Memory Network
abstract
Recently, memory-based methods have exhibited remarkable performance in Video Object Segmentation (VOS) by employing non-local pixel-wise matching between the query and memory. Nevertheless, these methods suffer from two limitations: 1) Non-local pixel-wise matching can result in the incorrect segmentation of background distractor objects, and 2) memory features with substantial temporal redundancy consume significant computing resources and reduce the inference speed. To address the limitations, we first propose a local attention mechanism to suppress background features, and we introduce a novel training framework based on contrast learning to ensure the network learns reliable and robust pixel-wise correspondence between query and memory. We adaptively determine whether to update the memory based on the variation of foreground objects. Next, we propose a dynamic memory bank, which utilizes a lightweight and differentiable soft modulation gate to determine the number of memory features to remove along the temporal dimension. This allows efficient and flexible management of memory features. Our network achieves competitive results (e.g., 92.1% on DAVIS 2016 val, 87.6%/81.3% on DAVIS 2017 val/test, 87.0% on YouTube-VOS 2018 val) compared with the state-of-the-art methods while maintaining a faster inference speed of 25+FPS. Moreover, our network demonstrates a favorable balance between performance and speed when dealing with the long-time video dataset.
Yadang Chen, Dingwei Zhang, Yuhui Zheng, Zhi-Xin Yang 0001, Enhua Wu, Haixing Zhao
IEEE Trans. Circuits Syst. Video Technol.4
2024 Multi-Sem Fusion: Multimodal Semantic Fusion for 3-D Object Detection
abstract
LIDAR and camera fusion techniques are promising for achieving 3D object detection in autonomous driving. Most multi-modal 3D object detection frameworks integrate semantic knowledge from 2D images into 3D LiDAR point clouds to enhance detection accuracy. Nevertheless, the restricted resolution of 2D feature maps impedes accurate re-projection and often induces a pronounced boundary-blurring effect, which is primarily attributed to erroneous semantic segmentation. To address these limitations, we present theMulti-Sem Fusion (MSF)framework, a versatile multi-modal fusion approach that employs 2D/3D semantic segmentation methods to generate parsing results for both modalities. Subsequently, the 2D semantic information undergoes re-projection into 3D point clouds utilizing calibration parameters. To tackle misalignment challenges between the 2D and 3D parsing results, we introduce an Adaptive Attention-based Fusion (AAF) module to fuse them by learning an adaptive fusion score. Then the point cloud with the fused semantic label is sent to the following 3D object detectors. Furthermore, we propose a Deep Feature Fusion (DFF) module to aggregate deep features at different levels to boost the final detection performance. The effectiveness of the framework has been verified on two public large-scale 3D object detection benchmarks by comparing them with different baselines. And the experimental results show that the proposed fusion strategies can significantly improve the detection performance compared to the methods using only point clouds and the methods using only 2D semantic information. Moreover, our approach seamlessly integrates as a plug-in within any detection framework.
Shaoqing Xu, Ziying Song, Sifen Wang, Zhi-Xin Yang 0001
IEEE Trans. Geosci. Remote. Sens.6
2024 Dynamic Focusing Network for Semisupervised Mechanical Fault Diagnosis of Rotating Machinery
abstract
The key components of the rotating machinery, such as gears and bearings, are prone to damage owing to long-term complex and harsh working situations. This study investigates the weight distribution of neural networks, and finds that the response of the network to the input is uneven, indicating that data-driven models tend to learn more information from certain local parts of the input. Based on this discovery, a novel attention mechanism, namely dynamic focusing, is proposed. The dynamic focusing mechanism highlights local information with important features to extract the discriminative features of key frequency bands. In addition, insufficient labeled data presents challenges for fault diagnosis in the practical. A semisupervised learning method based on mutual information is proposed to solve this problem. The effectiveness of the proposed method is verified by the Case Western Reserve University public dataset as well as the Gearbox Dynamic Simulator dataset obtained in our laboratory. The experimental results show that the proposed method has considerable advantages compared to existing deep learning methods, with test accuracy ranging from 95.31% to 100%.
Hao Chen 0099, Xianbo Wang, Jiaming Li 0019, Zhi-Xin Yang 0001
IEEE Trans. Ind. Informatics4
2024 Wind Turbine Fault Diagnosis for Class-Imbalance and Small-Size Data Based on Stacked Capsule Autoencoder
abstract
Wind power is of strategic importance for reducing carbon dioxide emissions, minimizing environmental pollution, and enhancing the sustainability of energy supply. Health monitoring of wind turbines is a crucial technology to ensure the quality of grid-connected power. Insufficient labeled data and class imbalance problems are two critical issues for intelligent fault diagnosis of wind turbines. In this article, an intelligent fault diagnosis method based on stacked capsule autoencoders is proposed to address the issues of inadequate labeled data and class imbalance. A prior knowledge-based convolution layer is applied to optimize the initialization of capsules, making it more conducive to learning spectral information. The pose representations of parts and objects can be improved, and a method for embedding spectral templates is proposed. The stacked capsule autoencoder in this study can learn partial templates unsupervised through likelihood estimation and establish the mapping between capsules and fault types. The experimental results, obtained from the CWRU dataset and a private dataset from a wind turbine drive-train simulation platform, demonstrate that the proposed method is robust to imbalanced and small-sized datasets. It can perform stable and effective unsupervised training by utilizing a sufficient amount of normal class data to expedite learning convergence.
Xianbo Wang, Hao Chen 0099, Jing Zhao 0010, Chonghui Song, Zhi-Xin Yang 0001, Pak-Kin Wong 0001
IEEE Trans. Ind. Informatics6
2024 Dual Branch Multi-Level Semantic Learning for Few-Shot Segmentation
abstract
Few-shot semantic segmentation aims to segment novel-class objects in a query image with only a few annotated examples in support images. Although progress has been made recently by combining prototype-based metric learning, existing methods still face two main challenges. First, various intra-class objects between the support and query images or semantically similar inter-class objects can seriously harm the segmentation performance due to their poor feature representations. Second, the latent novel classes are treated as the background in most methods, leading to a learning bias, whereby these novel classes are difficult to correctly segment as foreground. To solve these problems, we propose a dual-branch learning method. The class-specific branch encourages representations of objects to be more distinguishable by increasing the inter-class distance while decreasing the intra-class distance. In parallel, the class-agnostic branch focuses on minimizing the foreground class feature distribution and maximizing the features between the foreground and background, thus increasing the generalizability to novel classes in the test stage. Furthermore, to obtain more representative features, pixel-level and prototype-level semantic learning are both involved in the two branches. The method is evaluated on PASCAL-5i1-shot, PASCAL-5i5-shot, COCO-20i1-shot, and COCO-20i5-shot, and extensive experiments show that our approach is effective for few-shot semantic segmentation despite its simplicity.
Yadang Chen, Ren Jiang, Yuhui Zheng, Bin Sheng 0001, Zhi-Xin Yang 0001, Enhua Wu
IEEE Trans. Image Process.5
2024 Efficient odd-even multigrid for pointwise incompressible fluid simulation on GPU
Luan Lyu, Wei Cao 0008, Xiaohua Ren, Enhua Wu, Zhi-Xin Yang 0001
Vis. Comput.5
2023 Robust and Efficient Memory Network for Video Object Segmentation
abstract
This paper proposes a Robust and Efficient Memory Network, referred to as REMN, for studying semi-supervised video object segmentation (VOS). Memory-based methods have recently achieved outstanding VOS performance by performing non-local pixel-wise matching between the query and memory. However, these methods have two limitations. 1) Non-local matching could cause distractor objects in the background to be incorrectly segmented. 2) Memory features with high temporal redundancy consume significant computing resources. For limitation 1, we introduce a local attention mechanism that tackles the background distraction by enhancing the features of foreground objects with the previous mask. For limitation 2, we first adaptively decide whether to update the memory features depending on the variation of foreground objects to reduce temporal redundancy. Second, we employ a dynamic memory bank, which uses a lightweight and differentiable soft modulation gate to decide how many memory features need to be removed in the temporal dimension. Experiments demonstrate that our REMN achieves state-of-the-art results on DAVIS 2017, with a $\mathcal{J}\& \mathcal{F}$ score of 86.3% and on YouTube-VOS 2018, with a $\mathcal{G}$ over mean of 85.5%. Furthermore, our network shows a high inference speed of 25+ FPS and uses relatively few computing resources.
Yadang Chen, Dingwei Zhang, Zhi-Xin Yang 0001, Enhua Wu
ICME3
2023 An energy constraint position-based dynamics with corrected SPH kernel
Wei Cao 0008, Luan Lyu, Zhi-Xin Yang 0001, Enhua Wu
Sci. China Inf. Sci.3
2023 Spatial constraint for efficient semi-supervised video object segmentation
Yadang Chen, Chuanjun Ji, Zhi-Xin Yang 0001, Enhua Wu
Comput. Vis. Image Underst.3
2023 Global video object segmentation with spatial constraint module
abstract
We present a lightweight and efficient semi-supervised video object segmentation network based on the space-time memory framework. To some extent, our method solves the two difficulties encountered in traditional video object segmentation: one is that the single frame calculation time is too long, and the other is that the current frame’s segmentation should use more information from past frames. The algorithm uses a global context (GC) module to achieve high-performance, real-time segmentation. The GC module can effectively integrate multi-frame image information without increased memory and can process each frame in real time. Moreover, the prediction mask of the previous frame is helpful for the segmentation of the current frame, so we input it into a spatial constraint module (SCM), which constrains the areas of segments in the current frame. The SCM effectively alleviates mismatching of similar targets yet consumes few additional resources. We added a refinement module to the decoder to improve boundary segmentation. Our model achieves state-of-the-art results on various datasets, scoring 80.1% on YouTube-VOS 2018 and a $${\cal J}{\rm{\& }}{\cal F}$$ score of 78.0% on DAVIS 2017, while taking 0.05 s per frame on the DAVIS 2016 validation dataset.
Yadang Chen, Duolin Wang, Zhi-Xin Yang 0001, Enhua Wu
Comput. Vis. Media4
2023 Video object segmentation through semantic visual words matching
Chuanyan Hao, Yadang Chen, Zhi-Xin Yang 0001, Enhua Wu
Multim. Tools Appl.4
2023 Fixed-Time and Fault-Tolerant Path-Following Control for Autonomous Vehicles With Unknown Parameters Subject to Prescribed Performance
abstract
With the consideration of actuator faults, including the unknown steering mechanism misalignments and motor traction losses, this article presents a fixed-time control protocol to follow reference paths and velocities for autonomous ground vehicles (AGVs) with preset performance constraints. To provide sufficient large boundaries for the initial states, the hyperbolic tangent function is employed to predefine the constraints with respect to the path-following and velocity control performance. Based on the homeomorphic mapping and barrier Lyapunov theorem, the fixed-time prescribed performance control (PPC) objective-integrated fault-tolerant scheme can be achieved for the controlled AGV. In comparison to three different fixed-time controllers without the fault-tolerant or PPC scheme, the hardware-in-the-loop (HIL) test results demonstrate that the proposed control protocol can always provide superior control performance for the AGV under various maneuvering conditions.
Zhongchao Liang, Zhongnan Wang, Jing Zhao 0010, Pak-Kin Wong 0001, Zhi-Xin Yang 0001, Zhengtao Ding
IEEE Trans. Syst. Man Cybern. Syst.5
2023 Spatio-temporal compression for semi-supervised video object segmentation
Chuanjun Ji, Yadang Chen, Zhi-Xin Yang 0001, Enhua Wu
Vis. Comput.3
2022 Editorial Notes: Emerging intelligent automation and optimisation methods for adaptive decision making
Carman K. M. Lee, K. K. H. Ng, Roger Jianxin Jiao, Zhi-Xin Yang 0001
Adv. Eng. Informatics4
2022 Fast target-aware learning for few-shot video object segmentation
Yadang Chen, Chuanyan Hao, Zhi-Xin Yang 0001, Enhua Wu
Sci. China Inf. Sci.3
2022 Meta-transfer-adjustment learning for few-shot learning
Yadang Chen, Zhi-Xin Yang 0001, Enhua Wu
J. Vis. Commun. Image Represent.3
2022 Improving deep learning on point cloud by maximizing mutual information across layers
Di Wang 0032, Lulu Tang, Xu Wang 0037, Luqing Luo, Zhi-Xin Yang 0001
Pattern Recognit.5
2022 Mutual Information Maximization Based Similarity Operation for 3D Point Cloud Completion Network
abstract
Reconstructing the incomplete point cloud into the complete and uniform one is a fundamental task in 3D point cloud processing. Some existing studies have explored the use of deep learning networks and representation learning to implement point cloud completion, but the completed shapes still appear unrealistic and non-uniform. To address this issue, we propose the idea of simulating the process of point cloud completion with local-to-global reasoning (LGR). Motivated to achieve the fine LGR, a novel mutual information (MI) maximization-based similarity operation is proposed, which realizes the reconstruction from incomplete point cloud to complete one by maximizing MI between global features and the prior features from the same 3D object. We adopt the Jenson-Shannon MI estimator to maximize the MI between global features and the priors in specific dimensions, which effectively increases the similarity of global and prior features. Compared with other existing works and similarity operations, the superior results indicate the efficacy of the proposed method and its advantages over existing ones in both the synthetic and real-world datasets. Our source code is available athttps://github.com/wendydidi/MISO-PCN.
Di Wang 0032, Lulu Tang, Zhi-Xin Yang 0001
IEEE Signal Process. Lett.4
2022 Improving Semantic Analysis on Point Clouds via Auxiliary Supervision of Local Geometric Priors
abstract
Existing deep learning algorithms for point cloud analysis mainly concern discovering semantic patterns from the global configuration of local geometries in a supervised learning manner. However, very few explore geometric properties revealing local surface manifolds embedded in 3-D Euclidean space to discriminate semantic classes or object parts as additional supervision signals. This article is the first attempt to propose a unique multitask geometric learning network to improve semantic analysis by auxiliary geometric learning with local shape properties, which can be either generated via physical computation from point clouds themselves as self-supervision signals or provided as privileged information. Owing to explicitly encoding local shape manifolds in favor of semantic analysis, the proposed geometric self-supervised and privileged learning algorithms can achieve superior performance to their backbone baselines and other state-of-the-art methods, which are verified in the experiments on the popular benchmarks.
Lulu Tang, Ke Chen 0004, Chaozheng Wu, Kui Jia, Zhi-Xin Yang 0001
IEEE Trans. Cybern.6
2022 Statistically Evolving Fuzzy Inference System for Non-Gaussian Noises
abstract
Non-Gaussian noises always exist in the nonlinear system, which usually lead to inconsistency and divergence of the regression and identification applications. The conventional evolving fuzzy systems (EFSs) in common sense have succeeded to conquer the uncertainties and external disturbance employing the specific variable structure characteristic. However, non-Gaussian noises would trigger the frequent changes of structure under the transient criteria, which severely degrades performance. Statistical criterion provides an informed choice of the strategies of the structure evolution, utilizing the approximation uncertainty as the observation of model sufficiency. The approximation uncertainty can be always decomposed into model uncertainty term and noise term, and is suitable for the non-Gaussian noise condition, especially relaxing the traditional Gaussian assumption. In this article, a novel incremental statistical evolving fuzzy inference system (SEFIS) is proposed, which has the capacity of updating the system parameters, and evolving the structure components to integrate new knowledge in the new process characteristic, system behavior, and operating conditions with non-Gaussian noises. The system generates a new rule based on the statistical model sufficiency which gives so insight into whether models are reliable and their approximations can be trusted. The nearest rule presents the inactive rule under the current data stream and further would be deleted without losing any information and accuracy of the subsequent trained models when the model sufficiency is satisfied. In our article, an adaptive maximum correntropy extend Kalman filter is derived to update the parameters of the evolving rules to cope with the non-Gaussian noises problems to further improve the robustness of parameter updating process. The parameter updating process shares an estimate of the uncertainty with the criteria of the structure evolving process to make the computation less of a burden dramatically. The simulation studies show that the proposed SEFIS has faster learning speed and is more accurate than the existing evolving fuzzy systems (EFSs) in the case of noise free and noisy conditions.
Zhao-Xu Yang, Hai-Jun Rong, Plamen Angelov 0001, Zhi-Xin Yang 0001
IEEE Trans. Fuzzy Syst.4
2022 Dynamic Scene's Laser Localization by NeuroIV-Based Moving Objects Detection and LiDAR Points Evaluation
abstract
Accurate localization is an important component for the vehicle’s autonomous navigation. The appearance of the moving objects may lead to feature-matching error with the map features, thereby causing a serious decline of localization accuracy. Neuromorphic vision sensor (NeuroIV) is a kind of dynamic vision sensor, with the properties of high temporal resolution, movement capture, and lightweight computation. In view of this, this research proposes to combine the NeuroIV and LIDAR points to acquire the static landmark features and robust navigation localization. However, as a younger and smaller research field compared to RGB computer vision, NeuroIV vision is rarely associated with the intelligent vehicle. For this purpose, we built a novel dataset recorded by NeuroIV sensor, and a state-of-the-art YOLO-small network is designed to detect the moving objects with the dataset. In order to completely deduct the whole dynamic zones, a sensors’ novel fusion model is built by the zones’ segmentation and matching, so the LIDAR’s static environment is obtained completely by the remained points. By evaluating different types of LIDAR points, the feature-matching error can be alleviated further, making the localization is more accurate. Together with qualitative and quantitative results, this work provides a moving objects’ detection improvement of 14.13 % mAP with the new NeuroIV dataset, and an obvious localization accuracy improvement with LIDAR points’ evaluation.
Zhi-Xin Yang 0001
IEEE Trans. Geosci. Remote. Sens.2
2022 Distributed Adaptive Consensus Protocol for Connected Vehicle Platoon With Heterogeneous Time-Varying Delays and Switching Topologies
abstract
This paper studies the distributed consensus protocol for the connected vehicle platoon with heterogeneous time-varying delays and switching topologies. A third-order dynamics model with powertrain inertial lag is proposed to characterize the node longitudinal dynamics of vehicles in platoon. A novel distributed adaptive consensus protocol considering the time-varying delays and the random switched inter-vehicular communication topologies is designed to stabilize the heterogeneous vehicle platoon in the presence of external disturbance. The delay-range-dependent approach is used to deal with the system heterogeneous time-varying delays by considering the characteristics of the heterogeneous platoon. Directed graphs are adopted to describe the accessible information flow among vehicles. The necessary and sufficient conditions for the unified closed-loop vehicle platoon system are derived by using matrix analysis and Lyapunov-Krasovskii approach. Numerical simulations demonstrate the proposed method is effective.
Guokuan Yu, Pak-Kin Wong 0001, Jing Zhao 0010, Xianbo Wang, Zhi-Xin Yang 0001
IEEE Trans. Intell. Transp. Syst.6
2021 PU-EVA: An Edge-Vector based Approximation Solution for Flexible-scale Point Cloud Upsampling
abstract
High-quality point clouds have practical significance for point-based rendering, semantic understanding, and surface reconstruction. Upsampling sparse, noisy and non-uniform point clouds for a denser and more regular approximation of target objects is a desirable but challenging task. Most existing methods duplicate point features for upsampling, constraining the upsampling scales at a fixed rate. In this work, the arbitrary point clouds upsampling rates are achieved via edge-vector based affine combinations, and a novel design of Edge-Vector based Approximation for Flexible-scale Point clouds Upsampling (PU-EVA) is proposed. The edge-vector based approximation encodes neighboring connectivity via affine combinations based on edge vectors, and restricts the approximation error within a second-order term of Taylor’s Expansion. Moreover, the EVA upsampling decouples the upsampling scales with network architecture, achieving the arbitrary upsampling rates in one-time training. Qualitative and quantitative evaluations demonstrate that the proposed PU-EVA outperforms the state-of-the-arts in terms of proximity-to-surface, distribution uniformity, and geometric details preservation.
Luqing Luo, Lulu Tang, Wanyi Zhou, Shizheng Wang, Zhi-Xin Yang 0001
ICCV5
2021 A systematic literature review on intelligent automation: Aligning concepts from theory, practice, and future perspectives
K. K. H. Ng, Chun-Hsien Chen, Carman K. M. Lee, Roger Jianxin Jiao, Zhi-Xin Yang 0001
Adv. Eng. Informatics5
2021 Automatic representation and detection of fault bearings in in-wheel motors under variable load conditions
Xianbo Wang, Luqing Luo, Lulu Tang, Zhi-Xin Yang 0001
Adv. Eng. Informatics4
2021 A new nonlocal means based framework for mixed noise removal
Jielin Jiang, Jian Yang 0003, Zhi-Xin Yang 0001, Yadang Chen, Lei Luo 0001
Neurocomputing4
2021 Affine particle-in-cell method for two-phase liquid simulation
abstract
The interaction of gas and liquid can produce many interesting phenomena, such as bubbles rising from the bottom of the liquid. The simulation of two-phase fluids is a challenging topic in computer graphics. To animate the interaction of a gas and liquid, MultiFLIP samples the two types of particles, and a Euler grid is used to track the interface of the liquid and gas. However, MultiFLIP uses the fluid implicit particle (FLIP) method to interpolate the velocities of particles into the Euler grid, which suffer from additional noise and instability. To solve the problem caused by fluid implicit particles (FLIP), we present a novel velocity transport technique for two individual particles based on the affine particle-in-cell (APIC) method. First, we design a weighed coupling method for interpolating the velocities of liquid and gas particles to the Euler grid such that we can apply the APIC method to the simulation of a two-phase fluid. Second, we introduce a narrowband method to our system because MultiFLIP is a time-consuming approach owing to the large number of particles. Experiments show that our method is well integrated with the APIC method and provides a visually credible two-phase fluid animation. The proposed method can successfully handle the simulation of a twophase fluid.
Luan Lyu, Wei Cao 0008, Enhua Wu, Zhi-Xin Yang 0001
Virtual Real. Intell. Hardw.4
2020 Fracture Patterns Design for Anisotropic Models with the Material Point Method
abstract
Abstract Physically plausible fracture animation is a challenging topic in computer graphics. Most of the existing approaches focus on the fracture of isotropic materials. We proposed a frame‐field method for the design of anisotropic brittle fracture patterns. In this case, the material anisotropy is determined by two parts: anisotropic elastic deformation and anisotropic damage mechanics. For the elastic deformation, we reformulate the constitutive model of hyperelastic materials to achieve anisotropy by adding additional energy density functions in particular directions. For the damage evolution, we propose an improved phase‐field fracture method to simulate the anisotropy by designing a deformation‐aware second‐order structural tensor. These two parts can present elastic anisotropy and fractured anisotropy independently, or they can be well coupled together to exhibit rich crack effects. To ensure the flexibility of simulation, we further introduce a frame‐field concept to assist in setting local anisotropy, similar to the fiber orientation of textiles. For the discretization of the deformable object, we adopt a novel Material Point Method(MPM) according to its fracture‐friendly nature. We also give some design criteria for anisotropic models through comparative analysis. Experiments show that our anisotropic method is able to be well integrated with the MPM scheme for simulating the dynamic fracture behavior of anisotropic materials.
Wei Cao 0008, Luan Lyu, Xiaohua Ren, Bob Zhang 0001, Zhi-Xin Yang 0001, Enhua Wu
Comput. Graph. Forum5
2020 Higher-order potentials for video object segmentation in bilateral space
Chuanyan Hao, Yadang Chen, Zhi-Xin Yang 0001, Enhua Wu
Neurocomputing3
2020 An improved solution for deformation simulation of nonorthotropic geometric models
abstract
Abstract Physically based deformation simulation has been studied for many years in computer graphics. In order to simulate more complex geometric models and better meet the designer's requirements, many anisotropic approaches have been proposed in recent years. However, most of the approaches focus on simulating orthotropic models. In comparison with orthotropic models, nonorthotropic ones allow the objects to have anisotropic behaviors along nonorthogonal directions. In this paper, we introduce an improved approach to simulate nonorthotropic geometric models under large deformation. The improvements are mainly twofold. First, a frame field is specified on a given undeformed object, that is, each point of the object is equipped with a frame. In each local frame, we construct three independent vectors and form a nonorthogonal coordinate. Second, we design the deformation properties along each axis in the local nonorthogonal coordinate to get a local constitutive model. The final nonorthotropic model is generated by transforming the designed model from local nonorthogonal coordinates to the global standard Cartesian coordinate. To improve the stability, we introduce a time‐varying method to simultaneously track the local coordinates reorientation by pushing forward the original frame field to the deformed frame field. Experiments show that the deformation simulation using the designed nonorthotropic models exhibits anisotropic behaviors along different directions and are more stable than previous methods.
Wei Cao 0008, Zhi-Xin Yang 0001, Xiaohua Ren, Luan Lyu, Bob Zhang 0001, Yanci Zhang, Enhua Wu
Comput. Animat. Virtual Worlds2
2020 Single and simultaneous fault diagnosis of gearbox via a semi-supervised and high-accuracy adversarial learning framework
Pengfei Liang 0005, Jun Wu 0012, Zhi-Xin Yang 0001, Jinxuan Zhu
Knowl. Based Syst.4
2020 Adaptive neural tracking control for automotive engine idle speed regulation using extreme learning machine
Pak-Kin Wong 0001, Chi-Man Vong, Zhi-Xin Yang 0001
Neural Comput. Appl.4
2020 A new learning paradigm for random vector functional-link network: RVFL+
Peng-Bo Zhang, Zhi-Xin Yang 0001
Neural Networks2
2020 Robust and Noise-Insensitive Recursive Maximum Correntropy-Based Evolving Fuzzy System
abstract
In this article, a novel recursive maximum correntropy-based evolving fuzzy system (RMCEFS) is proposed. The proposed system has the capability of reorganizing the structure and adapting itself in a dynamically changing environment with non-Gaussian noises. The system generates a new rule based on the correntropy criterion which represents a robust nonlinear similarity measure between two random variables and avoids recruiting the noises as the rules. Maximizing the cross-correntropy between the system output and the desired response leads to the maximum correntropy criterion for system self-adaptation. In our article, a recursive solution of the maximum correntropy criterion is derived to update the parameters of the evolving rules. This avoids the convergence problem produced by the learning size in the gradient-based learning. Also, the steady-state convergence performance of the proposed RMCEFS is studied, where the analytical solutions of the steady-state excess mean square error for the Gaussian noise and non-Gaussian noises are derived. The simulation studies show that the proposed RMCEFS using the recursive maximum correntropy converges much faster and is more accurate than the existing evolving fuzzy systems in the case of noise-free and noisy conditions.
Hai-Jun Rong, Zhi-Xin Yang 0001, Pak-Kin Wong 0001
IEEE Trans. Fuzzy Syst.2
2019 Multilayer one-class extreme learning machine
Haozhen Dai, Jiuwen Cao, Tianlei Wang, Muqing Deng, Zhi-Xin Yang 0001
Neural Networks5
2019 A Novel Prognostic Approach for RUL Estimation With Evolving Joint Prediction of Continuous and Discrete States
abstract
In this paper, we propose a novel prognostic approach for remaining useful life estimation (RUL) with evolving joint prediction of continuous and discrete states which represent the signals and health states of systems respectively. The predictors are built with evolving capability of adapting structures and parameters online to capture the dynamic characteristics of systems during runtime. Moreover, the discrete states can be determined dynamically during the construction of the predictors for systems operating under different environments. In the testing phase, the optimum predictor for predicting continuous and discrete states jointly is chosen under the error and distance criteria. The RULs are estimated conveniently once the predicted signals fall into failure mode based on a distance metric. In order to validate the performance of the proposed approach, the widely used turbofan engine datasets are taken into consideration. Experimental results demonstrate the reasonability and superiority of the proposed approach compared to other approaches.
Rong-Jing Bao, Hai-Jun Rong, Zhi-Xin Yang 0001, Badong Chen
IEEE Trans. Ind. Informatics3
2018 Biorthogonal Wavelet Surface Reconstruction Using Partial Integrations
abstract
Abstract We introduce a new biorthogonal wavelet approach to creating a water‐tight surface defined by an implicit function, from a finite set of oriented points. Our approach aims at addressing problems with previous wavelet methods which are not resilient to missing or nonuniformly sampled data. To address the problems, our approach has two key elements. First, by applying a three‐dimensional partial integration, we derive a new integral formula to compute the wavelet coefficients without requiring the implicit function to be an indicator function. It can be shown that the previously used formula is a special case of our formula when the integrated function is an indicator function. Second, a simple yet general method is proposed to construct smooth wavelets with small support. With our method, a family of wavelets can be constructed with the same support size as previously used wavelets while having one more degree of continuity. Experiments show that our approach can robustly produce results comparable to those produced by the Fourier and Poisson methods, regardless of the input data being noisy, missing or nonuniform. Moreover, our approach does not need to compute global integrals or solve large linear systems.
Xiaohua Ren, Luan Lyu, Xiaowei He 0004, Wei Cao 0008, Zhi-Xin Yang 0001, Bin Sheng 0001, Yanci Zhang, Enhua Wu
Comput. Graph. Forum5
2018 A Novel AdaBoost Framework With Robust Threshold and Structural Optimization
abstract
The AdaBoost algorithm is a popular ensemble method that combines several weak learners to boost generalization performance. However, conventional AdaBoost.RT algorithms suffer from the limitation that the threshold value must be manually specified rather than chosen through a self-adaptive mechanism, which cannot guarantee a result in an optimal model for general cases. In this paper, we present a generic AdaBoost framework with robust threshold mechanism and structural optimization on regression problems. The error statistics of each weak learner on one given problem dataset is utilized to automate the choice of the optimal cut-off threshold value. In addition, a special single-layer neural network is employed to provide a second opportunity to further adjust the structure and strength the adaption capability of the AdaBoost regression model. Moreover, to consolidate the theoretical foundation of AdaBoost algorithms, we are the first to conduct a rigorous and comprehensive theoretical analysis on the proposed approach. We prove that the general bound on the empirical error with a fraction of training examples is always within a limited soft margin, which indicates that our novel algorithm can avoid over-fitting. We further analyze the bounds on the generalization error directly under probably approximately correct learning. The extensive experimental verifications on the UCI benchmarks have demonstrated that the performance of the proposed method is superior to other state-of-the-art ensemble and single learning algorithms. Furthermore, a real-world indoor positioning application has also revealed that the proposed method has higher positioning accuracy and faster speed.
Peng-Bo Zhang, Zhi-Xin Yang 0001
IEEE Trans. Cybern.2
2018 Single and Simultaneous Fault Diagnosis With Application to a Multistage Gearbox: A Versatile Dual-ELM Network Approach
abstract
High-precision fault diagnosis is vital for widely used multistage gearbox systems. Intelligent monitoring is difficult due to the fuzzy boundaries and a variety of unseen single or simultaneous faults of such complex machinery. To solve this problem, local mean decomposition is applied to extract features effectively from the original nonstationary and nonlinear vibration signals. By exploiting the diverse functionalities of extreme learning machines (ELM) in both regression and classification, a novel dual-ELM network is proposed, in which one ELM is employed to count the number of faults and the other is used to identify the specific single- or simultaneous-fault scenarios. The proposed dual-ELM-based multilabel classifier does not rely on an empirically specified threshold. Thus, it is more self-adaptive than the existing probabilistic-based classifiers. In addition, by inheriting the advantages of the original ELM, the dual-ELMs do not require iterative fine-tuning of parameters. Finally, the training speed of the dual-ELMs is much faster than other combinations of the existing classifiers. Experimental results under various loading conditions show that the proposed dual-ELM-based fault diagnostic framework is versatile at detecting single and simultaneous faults accurately and quickly.
Zhi-Xin Yang 0001, Xianbo Wang, Pak-Kin Wong 0001
IEEE Trans. Ind. Informatics1
2016 ELM Meets RAE-ELM: A hybrid intelligent model for multiple fault diagnosis and remaining useful life predication of rotating machinery
abstract
Reliable fault diagnosis and potential remaining useful life (RUL) predication before the occurrence of fatal failure in machinery is critical for improving productivity and reducing maintenance cost. However, the existing physics heuristics and neural networks based methods face difficulties to treat such two issues simultaneously. This paper proposes a novel Network of Extreme Learning Machines (N-ELM) framework, which is a hybrid model of classification and regression for multiple faults diagnosis and RUL predication. The N-ELM consists of a set of ELMs as the nodes of the learning network, which forms a “generalized” structure with fault detection and RUL forecasting functions. By exploiting the advantages of ELM superb efficiency in regression, the error statistics based robust AdaBoost.RT based ELM framework (RAE-ELM) with self-adaptive threshold mechanism is applied to improve the accuracy of RUL predication. The uniform network of multi-functioning ELMs enable classify fault types and predicate their corresponding RUL concurrently with outperformed accuracy and efficiency. The superior performance of the proposed hybrid framework and supporting techniques are validated using vibration monitoring dataset collected from rotating machinery in the field.
Zhi-Xin Yang 0001, Peng-Bo Zhang
IJCNN1
2016 Sparse Bayesian extreme learning committee machine for engine simultaneous fault diagnosis
Pak-Kin Wong 0001, Jianhua Zhong, Zhi-Xin Yang 0001, Chi-Man Vong
Neurocomputing3
2016 RFID-enabled indoor positioning method for a real-time manufacturing execution system using OS-ELM
Zhi-Xin Yang 0001
Neurocomputing1
2016 Fast detection of impact location using kernel extreme learning machine
Heming Fu, Chi-Man Vong, Pak-Kin Wong 0001, Zhi-Xin Yang 0001
Neural Comput. Appl.4
2014 Real-time fault diagnosis for gas turbine generator systems using extreme learning machine
Pak-Kin Wong 0001, Zhi-Xin Yang 0001, Chi-Man Vong, Jianhua Zhong
Neurocomputing2
2012 A systemic point-cloud de-noising and smoothing method for 3D shape reuse
abstract
3D shape reuse, as an effective way to carry out innovative design, requires a digital model database where the entities are accurate and sufficient representations of objects in the real world. 3D scanning is a prevailing tool to quickly convert physical models into virtual ones. However, the scanned models without post-processing could not be used directly due to environment noise and accuracy limitation in terms of discrete sampling property in scanning. This paper introduces a systemic point-cloud de-noising and mesh smoothing method to handle this issue. The model de-noising and regularity is based on k-means clustering, and mesh smoothing module is an improved mean approach which processes the discrete data in the regular order. Case study will be given to verify the smoothing effectiveness. The proposed method could facilitate the construction of model database for design reuse, and could be output to downstream applications such as shape adaptive deformation, and shape searching.
Zhi-Xin Yang 0001, Difu Xiao
ICARCV1