Jianxu Mao

dblp:93/10861 · DBLP profile ↗
← Back
43ranked-venue papers
1as first author
38since 2021 · last 2026
0000-0003-2267-5415ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 17 · 1 first-author · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 16 · 14 since 2021Artificial intelligence and machine learning · 9 · 7 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Distilling Object Detectors via Monte Carlo Dropout
abstract
Knowledge distillation (KD) has become a fundamental technique for model compression in object detection tasks. The data noise and training randomness may cause the knowledge of the teacher model to be unreliable, referred to as knowledge uncertainty. Existing methods neglect this uncertainty, potentially hindering the student's capacity to capture and understand latent "dark knowledge". In this work, we introduce a novel strategy that explicitly incorporates knowledge uncertainty, named Uncertainty-Driven Knowledge Extraction and Transfer (UET). Given the unknown, high-dimensional nature of the knowledge distribution, we employ Monte Carlo dropout to effectively estimate the teacher's uncertainty. Leveraging information theory, we combine uncertainty with deterministic knowledge, enabling the student to benefit from both precision and diversity. UET is a plug-and-play method that integrates seamlessly with existing distillation techniques. We validate our approach through comprehensive experiments across various distillation strategies, detectors, and backbones. Specifically, UET achieves state-of-the-art results, with a ResNet50-based GFL detector obtaining 44.1% mAP on the COCO dataset-surpassing baseline performance by 3.9%.
Junfei Yi, Hui Zhang 0023, Jianxu Mao, Tengfei Liu 0005, Mingjie Li 0006, Sihao Lin, Hanyu Gu, Zhihui Li 0001, Xiaojun Chang, Yaonan Wang 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
2026 High-Precision Multi-Instance Registration for Stacked Objects in Bin-Picking Scenes
abstract
In industrial bin-picking, robotic systems must estimate the poses of multiple object instances, where accurate pose estimation is essential for reliable downstream manipulation and grasping. Most existing multi-instance registration methods primarily establish point correspondences based on local features to alleviate the challenges posed by occlusion and clutter. However, local features are easily disturbed by neighboring instances and lack global context, leading to unreliable correspondences and degraded registration accuracy. In addition, the absence of rotational invariance further reduces correspondence accuracy in scenes with stacked instances and highly varying object orientations. To address these challenges, we present a one-stage multi-instance point cloud registration framework for stacked-object scenes. Our framework incorporates a rotation-invariant operator to enhance the robustness of feature representations under arbitrary orientations. Then, we propose a Center-Aware Res-Masked Transformer module, which incorporates an object center embedding to enrich global instance-level context and a center-aware residual mask prediction module to balance weight distribution across objects of varying sizes during training. Extensive experiments on the challenging ROBI dataset demonstrate that our method outperforms the competitive baseline MIRETR by more than 10% in mean precision, highlighting its effectiveness in complex bin-picking scenes. Furthermore, evaluations on the unstacked Scan2CAD dataset confirm the generalizability of the proposed framework across different application scenarios.
Jiawen Zhao, Qing Zhu 0003, Yaonan Wang 0001, Weixing Peng, Jianxu Mao, Min Liu 0008, Xuebing Liu, Hui Zhang 0023
IEEE Trans. Circuits Syst. Video Technol.5
2026 MGLD-TLNet: Multigeometric and Long-Distance Representation Network for Transmission Line Inspection
abstract
Effective transmission line (TL) inspection in complex corridor environments is essential for ensuring reliable power delivery. This work presents a 3D-based perception method for this task. The proposed method is designed by considering two key characteristics of TL inspection. First, the point cloud data are sparse and class distributions are highly imbalanced, which weakens the signals from thin conductors and tower components. To address this issue, we model long-range spatial relations along the corridor to mitigate data sparsity and imbalance. Second, strong structural correlations exist between conductors and towers, which can be leveraged to improve perception performance. To exploit this property, we construct a unified 3-D representation that jointly models towers, conductors, and vegetation, while fusing Cartesian and polar geometries through geometry-aware alignment. Experiments on real-world corridor datasets demonstrate that the proposed method, termed multigeometric and long-distance TL perception Network (MGLD-TLNet), consistently improves stability and accuracy under conditions of sparsity, occlusion, and complex environmental interactions.
Hui Zhang 0023, Kaining Zhang, Baheti Biekezat, Hang Zhong, Junfei Yi, Jianxu Mao, Yaonan Wang 0001
IEEE Trans. Cybern.7
2026 Investigating a Unified 3-D Object Detection Method for Different Multibeam LiDAR
Ziming Tao, Jianxu Mao, Yaonan Wang 0001, Caiping Liu, Junfei Yi, Zhenyu He 0015, Xiaojun Chang, Hui Zhang 0023
IEEE Trans. Ind. Informatics2
2026 SAF: A Structure-Aware Framework for Radial Ice Thickness Detection on Overhead Transmission Lines
abstract
Ice thickness estimation on overhead transmission lines (OHTL) is essential for mitigating icing-induced mechanical failures and ensuring safe grid operation. To address the challenges of detecting radial ice thickness in complex power line corridors, particularly geometric fragmentation of slender conductors and semantic ambiguity near occluded boundaries, this work proposes a structure-aware framework (SAF) based on 3-D point cloud segmentation and geometry-guided modeling. SAF introduces a structure-aware segmentation network, which integrates a cross-level spatial encoding module to preserve geometric continuity and a partition-aware loss to improve boundary localization under vegetation or tower occlusion. Building on accurate segmentation, a geometry-guided module performs centerline fitting and cross-sectional reconstruction to infer slice-level ice thickness. To support evaluation, a large-scale uncrewed aerial vehicle (UAV)-based point cloud dataset covering 32 OHTL is constructed, including six lines with ground-truth ice labels. Experimental results demonstrate that SAF achieves robust and accurate ice estimation across varied voltage levels and terrains, supporting its practical application in intelligent transmission line inspection and icing risk prevention.
Hui Zhang 0023, Youyuan Tang, Yihong Cao, Kaining Zhang, Yunkang Cao, Tongzhi Niu, Jianxu Mao, Yaonan Wang 0001
IEEE Trans. Ind. Informatics8
2026 You Can Only Tune Normalization: A Simple and Effective Approach to Parameter-Efficient Fine-Tuning
abstract
To tackle the issue of excessive parameter volumes during fine-tuning of large-scale pre-trained models with full parameters, Parameter-Efficient Fine-Tuning (PEFT) methods have been introduced. The core concept involves freezing the backbone network of the model and updating only a small subset of parameters. This strategy not only decreases the number of parameters needed for training but also delivers performance comparable to Full-Tuning, even surpassing it on certain datasets. However, most popular PEFT methods introduce extra parameters or modules for fine-tuning, which come with inherent limitations. In response, we propose a straightforward and efficient PEFT method called You Can Only Tune Normalization (YONO). YONO focuses solely on tuning the normalization layer and the final classification layer of the model. This method avoids adding extra modules, making it easily applicable to any model without causing inference delays. We extensively tested YONO on 28 benchmark datasets, and the results indicate that it requires significantly fewer parameters compared to other advanced PEFT methods. Additionally, we validated YONO’s efficiency and generalizability across various vision models. Finally, we further explore the essence of PEFT methods, whether they learn new knowledge or expose the capabilities that a model has already learned. Our findings suggest that YONO is more sensitive to improvements in dataset quality, making it a promising candidate for future scaling to larger models.
Lingyun Huang, Jianxu Mao, Junfei Yi, Ziming Tao, Ziyang Peng, Wei He 0001, Rui Liu 0028, Yaonan Wang 0001
ACM Trans. Intell. Syst. Technol.2
2026 Image-Quality-Guided Consistency-Alignment Network for Fault Diagnosis
abstract
Signal-to-image transformation has been widely used in mechanical fault diagnosis to provide unified visual representations for downstream diagnostic models. However, the resulting diagnostic images can exhibit substantial quality variations caused by measurement noise, sensor degradation, and partial signal loss. Most existing methods either discard low-quality samples or implicitly assume equal reliability across samples, leading to information loss or degraded robustness. This paper proposes an image-quality-guided consistency-alignment network that explicitly estimates sample quality and leverages degraded data during training. First, vibration signals are decomposed into sub-bands, and energy-ratio criteria are used to select informative components for reconstruction. The reconstructed signals are subsequently fused into fixed-layout two-dimensional images via a sector-allocation strategy. Next, a diagnosis-aware quality assessment module assigns a quality score to each image to guide training. Low-quality samples are further regularized via cross-quality feature alignment using a quality-weighted supervised pairwise loss, encouraging them to align with high-quality counterparts from the same class. Finally, a Swin Transformer backbone performs classification. Experiments on the proprietary RGFD dataset and the public SEUB dataset demonstrate consistent gains under controlled mixed-quality settings, with accuracies of up to 99.4% and 98.7%, respectively, while additional evaluations under reproducible degradations further confirm the robustness of the proposed quality-aware learning mechanism.
Zhuowei Li 0011, Jianxu Mao, Yaonan Wang 0001, Junfei Yi, Caiping Liu, Hui Zhang 0023
IEEE Trans. Reliab.2
2025 CVPT: Cross Visual Prompt Tuning
Lingyun Huang, Jianxu Mao, Junfei Yi, Ziming Tao, Yaonan Wang 0001
ICCV2
2025 Wavelet-Based Distillation with Structured Frequency Alignment
Pengyu Lu, Junfei Yi, Jianxu Mao, Junlong Yu, Shuohao Xiao, Zhenyu He 0015, Yaonan Wang 0001
ICIG (2)3
2025 Point Density Fusion for Multimodal 3D Object Detection
Ziyang Peng, Jianxu Mao, Wei He 0001, Caiping Liu, Zhenyu He 0015, Ziming Tao, Yaonan Wang 0001
ICIG (2)2
2025 FFTA-Net: A Frequency-Domain Fusion and Temporal Alignment Network for Transmission Line Defect Detection
Jianxu Mao, Yaonan Wang 0001, Junlong Yu, Junfei Yi, Zhenyu He 0015, Ziming Tao, Hui Zhang 0023
ICIG (2)2
2025 Foreground-Aware Enhancement-Based Multimodal 3D Object Detection
abstract
LiDAR is one of the most widely used 3D detection sensors in applications such as autonomous driving and unmanned inspection. However, when uniformly sampling the entire scene to generate point cloud data, the number of foreground points reflected by target objects is often significantly lower than that of the background points, which have a larger coverage area. This imbalance poses considerable challenges to the performance of object detection models, especially in detecting small or distant objects. To overcome this challenge, this paper presents a Foreground-Aware Enhancement-based Multimodal 3D Object Detection Method (PFA), which effectively mitigates the low detection accuracy of small and distant objects caused by insufficient foreground points. The proposed method incorporates a Foreground-Aware Enhancement Module (FAEM) and a Region-Focused Attention Module (RFAM). The FAEM module enhances the model’s focus on foreground regions, while the RFAM module strengthens multimodal fused features. Together, these components significantly improve the detection accuracy of small and distant objects. Experimental results on the KITTI dataset demonstrate that the proposed method achieves 3D detection accuracies of 84.65%, 59.56%, and 71.48% for cars, pedestrians, and cyclists, respectively, under the hard evaluation level. Furthermore, the model also shows significant advantages on the KITTI public test set and validation set for both easy and moderate samples, fully validating its effectiveness and generalizability in enhancing multimodal 3D object detection accuracy.
Ziyang Peng, Wei He 0001, Jianxu Mao, Ziming Tao, Junfei Yi, Yaonan Wang 0001
IJCNN3
2025 MQKIN: Manufacturing Quality Knowledge-Driven Interpretable Fault Diagnosis Network for Robotic Grinding Equipment
abstract
Despite of the fast development of deep learning networks, its inexplicability poses a low credibility challenge for fault diagnosis methods based on them. This article proposes an interpretable fault diagnosis network for robotic grinding equipment driven by the knowledge of grinding process. The network consists of a vibration imaging module, a grinding quality quantification module, an auxiliary learning module, and a fault diagnosis module. In addition, the developed synchronous algorithm is embedded in the network to update the weights and coefficient matrix combinations, which can reconstruct the vibration images from the wavelet domain to be recognized by convolution kernels more easily. Besides, through auxiliary learning, the grinding knowledge can be learned by the network to endow the results with interpretability. Finally, the comparative experiment shows that the proposed method has a maximum accuracy of 4.25%, 7.25%, 2.25%, 3.25%, and 1.25% higher than the five popular state-of-the-art fault diagnosis algorithms, respectively; the ablation experiment shows that the proposed algorithm has a maximum accuracy of 6.25%, 12.5%, and 16.25% higher than each sub-algorithm, respectively; by analyzing the learned weights of the network, it can be concluded that the proposed network has successfully learned the potential features of grinding knowledge with satisfactory performance.Note to Practitioners—In the autonomous robotic grinding system, the reliability of robotic grinding equipment is a prerequisite for ensuring high-quality grinding of thin-walled parts for large equipment such as aircraft, high-speed rail, and ships. In robotic grinding manufacturing, the information of grinding quality is the most direct feedback of valuable information from the manufacturing system. Based on the causal relationship between the fault status of grinding equipment and the grinding quality of workpiece, this article develops an interpretable fault diagnosis network incorporated with the grinding knowledge, which treats grinding quality information as physical knowledge and provides credibility to the fault diagnosis results. Therefore, this article improves the interpretability of the fault diagnosis model in practical applications through the credibility of data and grinding knowledge, which improves the applicable potential of fault diagnosis methods based on deep learning in industry.
Jianxu Mao, Yaonan Wang 0001, Zhe Li 0050, Xudong Wang 0008, Shaoyuan Wang
IEEE Trans Autom. Sci. Eng.2
2025 Prototype, Modeling, and Control of Aerial Robots With Physical Interaction: A Review
abstract
This article aims to investigate the research achievements related to aerial robots with physical interaction. Various morphologies of aerial physical interaction (APhI) robot prototypes with fixed wing, flapping wing, single main rotor, conventional underactuated multirotor, fully actuated multirotor, even deformed multirotor, and multiple platforms are reviewed for different APhI tasks associated with momentary, loose, and strong interaction coupling. This review also covers APhI robot rigid dynamics and robot-environment coupled interaction dynamics modeling methods, interaction wrench measurement/estimation, decoupled and coupled control, active aerial interaction control, and task-constrained planning approaches. Finally, future development directions and prospects are initially anticipated for aerial robots with physical interaction.Note to Practitioners—Aerial physical interaction (APhI) has been a hot topic in the field of aerial robots in recent years, which is a reflection of the advanced capabilities of aerial robots. However, APhI robots face challenges such as difficulty in flight stability and weak adaptability to dynamic environments while exerting active influence on environments. Under this background, this review aims to offer a reference for researchers and practitioners engaged in the related field from the aspects of system design, modeling, control, and task-constrained planning, which hopes to help them apply APhI robots to polar scientific expeditions, complex environment sampling, infrastructure inspection and maintenance, and other application areas. Further, this review also highlights the design idea of rigid-soft integrated APhI robots from the perspective of design-mechanism-performance to enhance interaction stability and safety.
Hang Zhong, Jiacheng Liang, Hui Zhang 0023, Jianxu Mao, Yaonan Wang 0001
IEEE Trans Autom. Sci. Eng.5
2025 LDFCDet: Boosting 3D Object Detectors With Low-High Level Feature Crosses Using Laplace Distribution
abstract
Highly accurate 3D object detection is critical for autonomous driving and robotic sensing system. However, some objects with few foreground points significantly affect the accuracy of 3D object detection. As the network depth increases, the low-level features of these objects are gradually lost, especially for the hard object. Due to this issue, current LiDAR-only based and multimodal methods often misclassify background as foreground. Therefore, how to leverage the low-level feature that contain information about these objects in the high layer of the network becomes the key to optimizing the issue. In this paper, we propose LDFCDet, a framework boosting 3D object detectors with low-high level feature crosses using Laplace distribution(LD). In our proposed method, we design a low-high level feature crosses module(LHFCM) to embed low-level feature into high-level feature in the deeper layer of the network, and use Laplace distribution to obtain a new low-high level feature that includes information about these objects with few foreground points. In addition, we propose a res-gated feature aggregation module(RGFAM) to fuse the mutli-scale features. Our approach is well-suited for both LiDAR-based and multimodal methods.We evaluate the LDFCDet on the widely used KITTI dataset, and our method outperforms almost current 3D object detection methods on the challenging KITTI test set. Moreover, we conducted comparative experiments on the ONCE dataset, and the results further demonstrate the effectiveness and superiority of our method.
Zhenyu He 0015, Jianxu Mao, Yaonan Wang 0001, Junlong Yu, Ziming Tao, Junfei Yi, Hui Zhang 0023, Shaoyuan Wang
IEEE Trans. Circuits Syst. Video Technol.2
2025 FMSD: Focal Multi-Scale Shape-Feature Distillation Network for Small Fasteners Detection in Electric Power Scene
abstract
In the electric power scene, fasteners play a pivotal role in securing and connecting electrical equipment, with small fastener detection (SFD) being crucial for ensuring operational stability. Despite the replacement of manual inspection methods by non-destructive techniques employing deep learning, these approaches often demand substantial computational resources and involve numerous parameters. While knowledge distillation (KD) can be a viable solution, existing KD methods may often fail to achieve satisfactory performance when dealing with small object presentation and little inter-class variability in SFD tasks. To alleviate this, we propose a Focal Multi-scale Shape-feature Distillation Network (FMSD) to achieve efficient and precise fastener detection in electric power scenarios. Specifically, we propose a novel Multi-Scale Shape-Aware Feature Aggregation module (MSFA) to augment the network's perception of object shape and scale during the KD process. Additionally, we propose a Contour-Guided Distillation (CGD) module to optimize the transfer of the extracted shape-sensitive knowledge between the teacher and student models. Through a series of experiments compared with existing state-of-the-art (SOTA) methods, our method demonstrates superior performance over existing SOTA techniques, both efficiently and effectively. Furthermore, validation on publicly available power scene datasets confirms the generalizability and adaptability of our proposed FMSD across various settings.
Junfei Yi, Jianxu Mao, Hui Zhang 0023, Mingjie Li 0006, Kai Zeng 0010, Mingtao Feng, Xiaojun Chang, Yaonan Wang 0001
IEEE Trans. Circuits Syst. Video Technol.2
2025 Toward Efficient Power Scene Detection via Topology-Preserved Knowledge Distillation
abstract
The power industry relies on efficient inspection systems to ensure stability and safety. While deep learning has advanced automated inspection, its reliance on custom modules for specific tasks can impact efficiency. Knowledge distillation (KD) offers a balanced solution, but the complex textures and structures of power equipment challenge conventional KD methods, which often fail to capture essential local semantic and topological relationships. To address this, we proposeTopNet, a novel topology-preserved KD framework for power scene detection tasks. Specifically, we model the teacher’s knowledge as a graph, where nodes encode local fine-grained features and edges capture global topological relationships. Based on this, we introduce node feature distillation and edge feature distillation to transfer local–global structural knowledge, which can enhance the student’s ability to perceive objects. Furthermore, we also introduce aggregated feature distillation to incorporate and transfer contextual semantic knowledge. Comprehensive experiments are conducted on two different benchmark datasets to demonstrate that TopNet achieves state-of-the-art detection performance with high efficiency, offering a robust solution for automated power equipment inspection.
Junfei Yi, Tengfei Liu 0005, Jianxu Mao, Yaonan Wang 0001, Hui Zhang 0023, He Xie, Hang Zhong, Xiaojun Chang
IEEE Trans. Ind. Informatics3
2025 Balancing Accuracy and Efficiency With a Multiscale Uncertainty-Aware Knowledge-Based Network for Transmission Line Inspection
abstract
Real-world transmission line inspections (RTLIs) ensure power stability and safety. Deep learning (DL) models have become prevalent approaches for performing RTLI tasks. However, the high computational demands and substantial parameter requirements of DL models limit their real-world applicability. This article introduces a novel approach, a multiscale uncertainty-aware knowledge-based network, which is designed to balance the accuracy and efficiency in RTLI tasks. Specifically, we propose an uncertainty-aware knowledge distillation method that incorporates pixel-level uncertainty into the knowledge transfer process, mitigating the impact of noisy knowledge derived from extra background information contained in ground truths. In addition, our method integrates a multiscale relationship distillation technique, thus enhancing the transfer of multiscale information between the teacher and student models. Consequently, RTLI tasks can be efficiently accomplished using the well-learned lightweight student model. Comprehensive experiments conducted on a real-world dataset collected via uncrewedaerial vehicles demonstrate the efficacy of our proposed approach in terms of achieving high detection accuracy with reduced computational costs.
Junfei Yi, Jianxu Mao, Hui Zhang 0023, Yurong Chen 0003, Tengfei Liu 0005, Kai Zeng 0010, He Xie, Yaonan Wang 0001
IEEE Trans. Ind. Informatics2
2025 Registration of Multiview Point Clouds With Unknown Overlap
abstract
Registration of multiview point clouds obtained from 3D scanners is a common method for 3D reconstruction. However, most existing registration methods are designed to handle point clouds with known overlap relationships that are ensured by external equipment (e.g., manipulators, turntables) or acquisition sequences, which limits the application range and increases the acquisition cost. To overcome these limitations, an unknown overlap registration (UOR) method for multiview point clouds is proposed, which can estimate overlap confidence, construct a connected graph, and remove outlier point clouds automatically. First, the overlap confidence between two point clouds is estimated by calculating the average nearest neighbor feature distance within the predicted overlap region. We then construct a minimal spanning tree based on the confidence levels and search for the central node to serve as the world coordinate. Finally, the Lie algebra-based SE(3)-sensitive perturbation scheme is introduced to solve the fine transformations, in which a robust weighting function is designed to weight point correspondences. Our method can find reliable connections among point clouds, and the proposed graph can be combined with different pairwise registration methods. The experimental results on both indoor and industrial datasets demonstrate the accuracy and effectiveness of our method.
Jiawen Zhao, Qing Zhu 0003, Yaonan Wang 0001, Weixing Peng, Hui Zhang 0023, Jianxu Mao
IEEE Trans. Multim.6
2025 Contrastive Learning Framework With Cross-Sensor Adaptive Signal Representation for Fault Diagnosis
abstract
Although multisource sensor (MS) signal-based mechanical fault diagnosis (MFD) can significantly improve the diagnostic performance, the existing methods often lack sufficient adaptability and generalization when retraining on single-sensor signals or inferring from partial sensor signals. Thus, a general two-stage signal representation contrastive learning fault diagnosis framework (T-SCF) is proposed to adapt the trained model to varying numbers of sensor signals. This framework enhances model robustness and data fusion by comparing sensor signal views, offering a new approach for information fusion, fault detection, and classification in MFD. In the first stage, an adaptive contrastive algorithm is proposed to generate contrastive samples (C-Ss) and contrastive labels (C-Ls) for MS signals. Then, a supervised contrastive loss (SCL) is designed to minimize the similarity between different fault MS signals while maximizing the similarity between identical ones. By designing a parallel encoder architecture, SCL enables it to merge contrasting the features of different sensor signals during training. This strategy preserves the time-domain dimension properties of different sensors during the training of the second-stage classifier, thereby improving the adaptability of the model to different sensor signals without affecting the global information. The effectiveness of the method was verified from multiple different evaluation dimensions using two public datasets and one self-built dataset.
Jianxu Mao, Yaonan Wang 0001, Zhe Li 0050, Hui Zhang 0023
IEEE Trans. Neural Networks Learn. Syst.2
2024 Structural and Textural-Aware Feature Extraction for Hyperspectral Image Classification
abstract
Feature extraction is a prevalent technique in hyperspectral remote sensing. Various tasks require this technique as a pre-processing step, including image classification, anomaly detection, image denoising, and so on. Edge-preserving filtering based methods have been extensively utilized for this purpose. However, these methods do not take the inherent structural and textural information into account, leading to poor performance in classifying hyperspectral images (HSIs). In this letter, a new structural and textural-aware feature extraction method is proposed that preserves the relevant structural information and removes useless textures. First, structural and textural-aware recursive filtering features (STRFs) are extracted along with an exponential form of windowed inherent variance (eWIV). Then, multi-scale STRFs are integrated by the principal component analysis (PCA) method to obtain more discriminative features (MSTRF). Finally, the fused features are fed into a pixel-wise classifier to obtain the final results. The main difference between the MSTRF method and other feature extraction methods is that the MSTRF method can make full use of the proposed eWIV map, which can help to properly characterize structure and texture in HSIs. Experimental results on several public data sets indicate that our method leads to state-of-the-art classification performance, especially in the presence of very small training set.
Ying Zhang 0063, Lianhui Liang, Jun Li 0009, Antonio Plaza, Xudong Kang, Jianxu Mao, Yaonan Wang 0001
IEEE Geosci. Remote. Sens. Lett.6
2024 Adaptive Force Tracking Impedance Control for Aerial Interaction in Uncertain Contact Environment Using Barrier Function
abstract
In this article, an adaptive force tracking impedance control strategy is investigated for an aerial manipulator in physical interaction with uncertain contact environments. Based on the modified target impedance model, an adaptive impedance control method is proposed to accomplish aerial interaction in uncertain environments while maintaining a stable contact force, wherein the environment parameters of location and stiffness are estimated online to generate a reference position trajectory. Then, in order to ensure the tracking performance of the aerial manipulator, a robust pose tracking controller is designed, including a barrier function-based position controller and an adaptive attitude controller. Both proposed position and attitude controllers can ensure finite-time convergence of the state variable without the priori boundary information of disturbances. In particular, the position state variable can converge to a predefined neighborhood of zero from any initial state, and the control gain is not overestimated. The stability of the proposed strategy is analyzed via Lyapunov tools. Simulations and real-world experiments are conducted to illustrate the feasibility and performance of the proposed control strategy.Note to Practitioners—The motivation of this article is to investigate an adaptive force tracking impedance control strategy for aerial physical interaction with uncertain contact environments. In the existing impedance control schemes for aerial manipulators, the environment parameter of location or stiffness is often required to be utilized in controller design. However, in practical cases, the environmental parameters are not known precisely. Thus, this article presents an adaptive impedance method to automatically generate the reference position trajectory and achieve a stable contact force. Additionally, the tracking performance of the aerial manipulator is inevitably subject to uncertainties and disturbances. To ensure tracking convergence, traditional robust controllers generally involve high control gains than the known upper bounds of the disturbances. The main disadvantage of those controllers is that the control gain is often overestimated when the disturbance decreases. To address this issue, a barrier function-based position controller is proposed for the aerial manipulator, where the priori boundary information of disturbances is not needed and the control gain is adaptively adjusted according to the amplitude of disturbances. The stability and convergence of the proposed strategy are analyzed mathematically, and the experiments using an aerial manipulator provide promising results.
Jiacheng Liang, Hang Zhong, Yaonan Wang 0001, Junhao Zeng, Jianxu Mao
IEEE Trans Autom. Sci. Eng.6
2024 Hybrid Force/Position Control of Multi-Mobile Manipulators for Cooperative Operation Without Force Measurements
abstract
In this paper, considering the difficulty of the interaction between the multi-mobile manipulators and the environment, the dynamics model of the mobile manipulator is analyzed, and a hybrid force/position control method based on the prescribed performance is proposed to improve the stability of the multi-mobile manipulators in the process of cooperative object transportation. Firstly, the dynamics model of the underdriven system of the multi-mobile manipulators is established by the Newton-Euler theorem. Then, the equivalent control theory is adopted for underdriven system, and a prescribed performance control method is proposed by considering the motion interference between the mobile manipulator and the high precision control of the manipulator. At the same time, an adaptive impedance control method is used to overcome internal and external disturbance during the cooperative transport of multi-mobile manipulators. The stability of the proposed method is analyzed through the Lyapunov stability theory. Finally, the effectiveness and superiority of the proposed scheme are verified through a simulation of multi-mobile manipulators collaborative object transportation.
Jianxu Mao, Haoran Tan, Yiming Jiang 0001, Yun Feng 0001, You Wu 0005, Yaonan Wang 0001
IEEE Trans. Circuits Syst. I Regul. Pap.2
2024 Category-Contextual Relation Encoding Network for Few-Shot Object Detection
abstract
Few-shot object detection (FSOD) has brought increasing academic interest by recognizing previously unseen novel classes with very limited well-labeled samples. However, most existing methods identify novel classes via some object-specific characteristics in the few provided samples rather than intrinsic inter-class relations between base and novel classes, which heavily degrades the detection performance on novel classes. Moreover, they cannot learn discriminative proposal representations to distinguish base and novel classes, and thus misclassify novel objects as confusable base classes. To tackle the above challenges, we develop a novel Category-contextual Relation Encoding Network (CRE-Net), which is an early attempt to reason inter-class context relationships for FSOD task. To be specific, we propose a novel category-contextual relation encoding mechanism to capture intrinsic inter-class relations between base and novel classes via knowledge aggregation from global category-contextual descriptors. It utilizes intrinsic inter-class contextual relations to adaptively refine the convolution kernel, thus encoding the local semantic context of query image with category-contextual relation as guidance. Furthermore, to explore discriminative representations for base and novel classes, we develop a scarcity-compensatory contrastive proposal loss by incorporating data scarcity of novel classes and proposal semantic consistency with high confidence. This loss could compact object instances from the same category to a tighter cluster, and enhance the space separability of different classes. Extensive experiments on Pascal VOC and COCO datasets verify the state-of-the-art detection performance of our CRE-Net model when compared with other baseline methods.
Ating Yin, Yaonan Wang 0001, Jianxu Mao, Hui Zhang 0023, Xiuyi Chen
IEEE Trans. Circuits Syst. Video Technol.3
2024 Deep Stereo Network With MRF-Based Cost Aggregation
abstract
Despite the remarkable progress made in learning-based stereo-matching algorithms, it is an open challenge for stereo-matching in disparity discontinuities and textureless regions. In this paper, we propose the deep Markov Random Field based cost aggregation network (DMCA-Net) for stereo matching, which is an end-to-end model-driven network architecture. This architecture introduces an efficient feature extraction network to extract richer textual and contextual feature information for stereo feature similarity representation at multi-stages and levels. Furthermore, with the aim of alleviating the edge-fattening phenomenon at disparity discontinuities and generating accurate disparities in textureless regions, we proposed the differentiable Markov Random Field model for cost aggregation, where the model’s data term utilizes image detail information, such as boundary and contour features, to guide matching cost aggregation, and the model’s smoothness term penalizes the adjacency similarity of the cost between the four-nearest neighboring pixel pairs to predict the disparity in textureless regions. The detailed experiment demonstrates that DMCA network achieves competitive performance on the SceneFlow, KITTI 2012, KITTI 2015, and Middlebury 2014 datasets.
Kai Zeng 0010, Hui Zhang 0023, Wei Wang 0025, Yaonan Wang 0001, Jianxu Mao
IEEE Trans. Circuits Syst. Video Technol.5
2024 Robust Variable Impedance Control for Aerial Compliant Interaction With Stability Guarantee
abstract
This article investigates a robust variable impedance control methodology for aerial manipulators to realize compliant and safe interaction tasks. Considering that the stability characteristics are generally overlooked in existing variable impedance controllers of the aerial manipulator, state-independent stability conditions are applied for time-varying impedance profiles to ensure the exponential stability of the desired variable impedance dynamics (DVID) as well as the boundedness of the state variables in the DVID. A command trajectory variable is introduced for converting the impedance control issue to a particular tracking issue, and then, a robust variable impedance controller based on the wrench estimator is designed to guarantee the exponential convergence of the translational states and impedance error of the aerial manipulator. The designed impedance controller is structurally simple and results in low implementation costs. Next, an improved attitude control approach with the command filter is developed for global flight attitude stability without any singularities or ambiguities, where the filter is introduced to avoid computing the derivative signals of the generalized force input. Finally, the effectiveness of the proposed control method is illustrated via numerical simulations and interaction experiments with different targets in real scenarios.
Jiacheng Liang, Yaonan Wang 0001, Hang Zhong, Hongwen Li, Jianxu Mao, Wei Wang 0025
IEEE Trans. Ind. Informatics6
2024 MSRN: Multilevel Spatial Refinement Network for Transmission Line Fastener Defect Detection
abstract
Transmission line (TL) fasteners play the role of connecting components in smart-grid transmission processes with abnormal TL fastener states, seriously impacting the power supply. Therefore, regular detection of TL fasteners is significant. However, the images taken by unmanned aerial vehicle have problems, such as small size and complex background, which bring great challenges to the existing object detection models. Based on this, this article proposes a multilevel spatial refinement network (MSRN), including an attention-guided receptive field enhanced feature pyramid network (ARFE-FPN) and a double refinement head (DR-Head). For the small target problem, ARFE-FPN first uses dilated convolution to expand the receptive field, and uses global average pooling to extract background activation values. Then, it performs channel weighting on TL fastener features under different fields of view. For the problem of complex background, DR-Head first constructs a semantic prediction task to realize the preseparation of foreground and background, and then combines the high-resolution feature map to further highlight the fastener features in the low-resolution feature map. Experiments on the TL fastener dataset show that MSRN has the best detection accuracy, and its AP can reach 92$\%$.
Jianxu Mao, Qingxian Liu, Yaonan Wang 0001, Weixing Peng, Junfei Yi, Ziming Tao, Hui Zhang 0023, Caiping Liu
IEEE Trans. Ind. Informatics1
2023 Deep Stereo Matching with Superpixel Based Feature and Cost
Kai Zeng 0010, Hui Zhang 0023, Wei Wang 0025, Yaonan Wang 0001, Jianxu Mao
PRCV (2)5
2023 Separable-programming based probabilistic-iteration and restriction-resolving correlation filter for robust real-time visual tracking
Baiheng Cao, Xuedong Wu, Jianxu Mao, Yaonan Wang 0001
Eng. Appl. Artif. Intell.3
2023 SSAPN: Spectral-Spatial Anomaly Perception Network for Unsupervised Vaccine Detection
abstract
Vaccines are the most significant and effective way to prevent disease and safeguard human health. However, it is easy to produce or mix foreign matters during the manufacturing process. Moreover, foreign matters are extremely faint that it is difficult to obtain images and detect them accurately. To tackle imaging challenges, in this article, we built a hyperspectral imaging system to construct a first-of-its-kind HSI dataset with pixel-level annotation for vaccine anomaly detection, where the vaccine comes from the actual pharmaceutical company. To address the problem of low detection accuracy, we propose a spectral–spatial anomaly perception module joint with an unsupervised autoencoder network (SSAPN), in which nonlinear features learned from the encoder are divided into nonoverlapping patches and mapped to efficiently encode spectral and spatial feature information. The spectral–spatial multilayer perceptrons (MLP) module consists of continuous and alternating spectral MLP with spatial MLP, which achieves spectral with spatial perception in the global receptive field, captures long-range dependencies, and extracts the most discriminative spectral–spatial features. Experimental results show that our SSAPN model outperforms other state-of-the-art anomaly detection methods in terms of both detection and generalization performance. This work will help speed up the production process in the vaccine pharmaceutical industry and ensure vaccine quality.
Ating Yin, Yaonan Wang 0001, Yurong Chen 0003, Kai Zeng 0010, Hui Zhang 0023, Jianxu Mao
IEEE Trans. Ind. Informatics6
2023 Deep Confidence Propagation Stereo Network
abstract
Stereo matching depth estimation based on rectified image pairs is of great importance to many computer vision tasks such as vehicle navigation and autonomous driving. Confidence measures are typically used to refine stereo matching results, which provides robustness and efficiency for disparity estimation. However, previous learning-based confidence methods for stereo matching usually use the middle results or composition as a post-processing step to refine the stereo matching results. This cannot be optimized end-to-end and the performance is limited by the quality of the tri-modal output. To handle this issue, in this paper, we pursue an end-to-end hierarchical architecture and propose a differentiable confidence propagation (DCP) model of a cost aggregation network for stereo matching. The DCP model is integrated into an end-to-end neural network hierarchical architecture to guide matching cost volume aggregation. More specifically, to better represent the similarity of left and right feature maps, we extract unary context feature maps with an effective attention mechanism for matching cost construction. Moreover, we aggregate the cost volume with the multiple stacked DCP cost aggregation (DCPCA) networks to generate a more reliable and finer cost volume. This network suppresses multi-level disparity maps. Each output disparity is supervised with different training weights to learn in a coarse-to-fine way. Our method outperforms previous methods on the Sceneflow dataset by achieving the$0.6735px$EPE error, achieving 1.53% D1-all metric of Non-occluded pixels regions and 0.72% Non-occluded pixels of$5px$metric on KITTI 2015 and 2012 dataset. Extensive experiments carried out on the KITTI Stereo benchmarks demonstrate that our DCPCA-Net can significantly minimize the trade-off between accuracy and efficiency for stereo matching.
Kai Zeng 0010, Yaonan Wang 0001, Wei Wang 0025, Hui Zhang 0023, Jianxu Mao, Qing Zhu 0003
IEEE Trans. Intell. Transp. Syst.5
2022 Review on the COVID-19 pandemic prevention and control system based on AI
Junfei Yi, Hui Zhang 0023, Jianxu Mao, Yurong Chen 0003, Hang Zhong, Yaonan Wang 0001
Eng. Appl. Artif. Intell.3
2022 Contour Structural Profiles: An Edge-Aware Feature Extractor for Hyperspectral Image Classification
abstract
Feature extraction provides an effective tool to classify hyperspectral images (HSIs). However, most hyperspectral feature extraction methods tend to yield an over-smoothed phenomenon, which leads to inconsistency between the homogeneous regions and the ground objects in the actual scene. To alleviate this problem, an edge-aware feature extractor called contour structural profiles (CSPs) is proposed to extract the discriminative features for hyperspectral images classification (HSIC). The proposed classification method comprises three components. First, the spectral dimension of the HSI is reduced with an averaging-based method. Then, an edge-aware total variation (TV) model is constructed to extract the contour structural profile, in which a learned contour probability map is served as one of the major cues in the feature extraction process. Next, multiscale structural profiles (MSSPs) are constructed using the edge-aware TV model with different parameters so as to fully characterize ground objects with different scales. Finally, the MSSPs are fused with a kernel principal component analysis (KPCA) followed by a spectral classifier to obtain the final classification map. Experimental results on several publicly available hyperspectral datasets illustrate that the proposed method obtains superior classification performance over several state-of-the-art classification approaches, especially when the number of training samples is insufficient.
Ying Zhang 0063, Puhong Duan, Jianxu Mao, Xudong Kang, Leyuan Fang, Pedram Ghamisi
IEEE Trans. Geosci. Remote. Sens.3
2022 Deep Stereo Matching With Hysteresis Attention and Supervised Cost Volume Construction
abstract
Stereo matching disparity prediction for rectified image pairs is of great importance to many vision tasks such as depth sensing and autonomous driving. Previous work on the end-to-end unary trained networks follows the pipeline of feature extraction, cost volume construction, matching cost aggregation, and disparity regression. In this paper, we propose a deep neural network architecture for stereo matching aiming at improving the first and second stages of the matching pipeline. Specifically, we show a network design inspired by hysteresis comparator in the circuit as our attention mechanism. Our attention module is multiple-block and generates an attentive feature directly from the input. The cost volume is constructed in a supervised way. We try to use data-driven to find a good balance between informativeness and compactness of extracted feature maps. The proposed approach is evaluated on several benchmark datasets. Experimental results demonstrate that our method outperforms previous methods on SceneFlow, KITTI 2012, and KITTI 2015 datasets.
Kai Zeng 0010, Yaonan Wang 0001, Jianxu Mao, Caiping Liu, Weixing Peng, Yin Yang 0004
IEEE Trans. Image Process.3
2022 Deep Progressive Fusion Stereo Network
abstract
Stereo matching depth estimation for rectified image pairs is of great importance to many compute vision tasks, specifically in autonomous driving. With the flourishing of convolution neural networks, responsible depth estimation of stereo matching with artificial intelligence is the most severe challenge for autonomous driving in recent years. Previous research on end-to-end trainable stereo matching networks has usually used cascading convolution blocks with down-sampling or pooling operations to extract the unary features required for matching cost construction. Such approaches lack a reconstruction stage for increasing feature map pixel-wise alignment and strength, factors which play an important role in representing the similarity between stereo image pairs. To address this issue, in this paper, we propose the progressive fusion stereo matching network (PFSM-Net). We exploit an encoder-decoder feature extraction network architecture for multi-stage and -scale dynamic feature extraction. Moreover, we propose a group-wise concatenation method to construct the cost volume, which provides a more efficient cost volume for cost aggregation. Furthermore, we propose the use of multi-scale cost aggregation networks with a progressive fusion strategy. The aggregated cost volume is progressively fused with the multi-stage and -scale cost volume as the size of the cost volume increases. Multi-stage and -scale outputs are supervised with and learned in a coarse-to-fine manner. Experimental results demonstrate that our method outperforms previous methods on the SceneFlow, KITTI 2012, and KITTI 2015 datasets.
Kai Zeng 0010, Yaonan Wang 0001, Qing Zhu 0003, Jianxu Mao, Hui Zhang 0023
IEEE Trans. Intell. Transp. Syst.4
2021 Semi-supervised Cloud Edge Collaborative Power Transmission Line Insulator Anomaly Detection Framework
Yanqing Yang, Jianxu Mao, Hui Zhang 0023, Yurong Chen 0003, Hang Zhong, Yaonan Wang 0001
ICIG (1)2
2021 Edge Guided Structure Extraction for Hyperspectral Image Classification
abstract
In this paper, a novel edge guided structure extraction method is proposed for hyperspectral images classification, which consists of the following steps: First, the spectral dimension of the hyperspectral image is reduced with an averaging-based method. Then, the structural features is extracted by an extended relative total variation (ERTV) inspired by a learned edge probability map which serves as one of the major cues in the structure extraction process. Finally, the extracted structural features are fed into SVM for classification. Experimental results on two publicly available hyperspectral data sets demonstrate the competitive performance over several state-of-the-art classification approaches.
Ying Zhang 0063, Puhong Duan, Xudong Kang, Jianxu Mao
IGARSS4
2021 Deep residual deconvolutional networks for defocus blur detection
abstract
Abstract Accurate defocus blur detection has instigated wide research interest for the last few years. However, it is still a meaningful yet challenging machine vision task, and most methods rely on prior knowledge. Convolutional neural networks have proved the huge success for different tasks within the computer vision, and machine learning flew. A simple yet effective method of defocus blur detection was proposed in this paper, which by applying the deep residual convolutional encoder‐decoder network. The aims of DRDN is to automatically generate pixel‐level predictions for defocus blur images, and reconstruct output detection results of the same size as the input, which by performing several deconvolution operations at multiple scales through the transposed convolution, and skip connection. Afterwards, we used the slide window detection strategy and traversed the input image with a certain stride. Experiments on challenging benchmarks of defocus blur detection show that our algorithm achieved state‐of‐the‐art performance, and powerfully balanced the detection accuracy, and detection time.
Kai Zeng 0010, Yaonan Wang 0001, Jianxu Mao, Xianen Zhou
IET Image Process.3
2020 Research on multi-focus image fusion algorithm based on total variation and quad-tree decomposition
Caiping Liu, Jianxu Mao
Multim. Tools Appl.3
2020 Autonomous mobile robot path planning in unknown dynamic environments using neural dynamics
Jiacheng Liang, Yaonan Wang 0001, Qi Pan, Jianhao Tan, Jianxu Mao
Soft Comput.6
2020 A Surface Defect Detection Framework for Glass Bottle Bottom Using Visual Attention Model and Wavelet Transform
abstract
Glass bottles must be thoroughly inspected before they are used for packaging. However, the vision inspection of bottle bottoms for defects remains a challenging task in quality control due to inaccurate localization, the difficulty in detecting defects in the texture region, and the intrinsically nonuniform brightness across the central panel. To overcome these problems, we propose a surface defect detection framework, which is composed of three main parts. First, a new localization method named entropy rate superpixel circle detection (ERSCD), which combines least-squares circle detection and entropy rate superpixel (ERS) with an improved randomized circle detection, is proposed to accurately obtain the region of interest (ROI) of the bottle bottom. Then, according to the structure-property, the ROI is divided into two measurement regions: central panel region and annular texture region. For the former, a defect detection method named frequency-tuned anisotropic diffusion super-pixel segmentation (FTADSP) that integrates frequency-tuned salient region detection (FT), anisotropic diffusion, and an improved superpixel segmentation is proposed to precisely detect the regions and boundaries of defects. For the latter, a defect detection strategy called wavelet transform multiscale filtering (WTMF) based on a wavelet transform and a multiscale filtering algorithm is proposed to reduce the influence of texture and to improve the robustness to localization error. The proposed framework is tested on four data sets obtained by our designed vision system. The experimental results demonstrate that our framework achieves the best performance compared with many traditional methods.
Xianen Zhou, Yaonan Wang 0001, Qing Zhu 0003, Jianxu Mao, Changyan Xiao, Xiao Lu 0002, Hui Zhang 0023
IEEE Trans. Ind. Informatics4
2019 A Local Metric for Defocus Blur Detection Based on CNN Feature Learning
abstract
Defocus blur detection is an important and challenging task in computer vision and digital imaging fields. Previous work on defocus blur detection has put a lot of effort into designing local sharpness metric maps. This paper presents a simple yet effective method to automatically obtain the local metric map for defocus blur detection, which based on the feature learning of multiple convolutional neural networks (ConvNets). The ConvNets automatically learn the most locally relevant features at the super-pixel level of the image in a supervised manner. By extracting convolution kernels from the trained neural network structures and processing it with principal component analysis, we can automatically obtain the local sharpness metric by reshaping the principal component vector. Meanwhile, an effective iterative updating mechanism is proposed to refine the defocus blur detection result from coarse to fine by exploiting the intrinsic peculiarity of the hyperbolic tangent function. The experimental results demonstrate that our proposed method consistently performed better than the previous state-of-the-art methods.
Kai Zeng 0010, Yaonan Wang 0001, Jianxu Mao, Junyang Liu, Weixing Peng, Nankai Chen
IEEE Trans. Image Process.3
2014 Adaptive motion/force control strategy for non-holonomic mobile manipulator robot using recurrent fuzzy wavelet neural networks
Yaonan Wang 0001, Thanglong Mai, Jianxu Mao
Eng. Appl. Artif. Intell.3