VLDB 2026 Research / reviewers in the wild / expert
Yaonan Wang 0001
dblp:90/548-1
· DBLP profile ↗
443ranked-venue papers
5as first author
361since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 191 · 4 first-author · 135 since 2021Applied, interdisciplinary, general and emerging computing · 129 · 1 first-author · 117 since 2021Graphics, computer vision, multimedia, augmented reality and games · 111 · 95 since 2021Systems, architecture and hardware · 26 · 23 since 2021Human-computer interaction and ubiquitous computing · 26 · 23 since 2021Security and privacy · 3 · 3 since 2021Computer networks · 2 · 1 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Mono3DVG-EnSD: Enhanced Spatial-aware and Dimension-decoupled Text Encoding for Monocular 3D Visual GroundingabstractMonocular 3D Visual Grounding (Mono3DVG) is an emerging task that locates 3D objects in RGB images using text descriptions with geometric cues. However, existing methods face two key limitations. Firstly, they often over-rely on high-certainty keywords that explicitly identify the target object while neglecting critical spatial descriptions. Secondly, generalized textual features contain both 2D and 3D descriptive information, thereby capturing an additional dimension of details compared to singular 2D or 3D visual features. This characteristic leads to cross-dimensional interference when refining visual features under text guidance. To overcome these challenges, we propose Mono3DVG-EnSD, a novel framework that integrates two key components: the CLIP-Guided Lexical Certainty Adapter (CLIP-LCA) and the Dimension-Decoupled Module (D2M). The CLIP-LCA dynamically masks high-certainty keywords while retaining low-certainty implicit spatial descriptions, thereby forcing the model to develop a deeper understanding of spatial relationships in captions for object localization. Meanwhile, the D2M decouples dimension-specific (2D/3D) textual features from generalized textual features to guide corresponding visual features at same dimension, which mitigates cross-dimensional interference by ensuring dimensionally-consistent cross-modal interactions. Through comprehensive comparisons and ablation studies on the Mono3DRefer dataset, our method achieves state-of-the-art (SOTA) performance across all metrics. Notably, it improves the challenging Far([email protected]) scenario by a significant +13.54%. Min Liu 0008, Zhaoyang Li 0011, Yuan Bian 0002, Erbo Zhai, Yaonan Wang 0001 |
AAAI | 7 |
| 2026 | HSFusion: hierarchical multi-scale feature fusion network based on state space model for infrared and visible image fusion
Huaying Cheng, Zhi Li 0087, Qin Wan 0001, Yanhui Xi, Shao Sheng Fan, Haotian Wu 0002, Yaonan Wang 0001 |
Expert Syst. Appl. | 8 |
| 2026 | Multiobjective harvester scheduling with splittable workloads: A diversity-guided iterated local search approach
Jing J. Liang, Caitong Yue, Yaonan Wang 0001 |
Expert Syst. Appl. | 4 |
| 2026 | One-shot motion talking head generation with audio-driven model
Weiliang Meng, Yaonan Wang 0001 |
Expert Syst. Appl. | 4 |
| 2026 | MarineSeg: A CNN-transformer hybrid architecture with feature voting decoder for robust semantic segmentation in USV-captured images
Qingyang Gu, Baoyuan Deng, Yunze He, Liang Cheng 0005, Yaonan Wang 0001 |
Neurocomputing | 6 |
| 2026 | SWS-YOLO: An energy-efficient spiking neural network for water-surface object detection
Yunze He, Baoyuan Deng, Liang Cheng 0005, Yaonan Wang 0001 |
Neurocomputing | 6 |
| 2026 | SCVI: A semi-coupled visible-infrared small object detection method based on multimodal proposal-level probability fusion strategy
Haozhi Xu, Xiaofang Yuan, Yaonan Wang 0001 |
Neurocomputing | 4 |
| 2026 | A two-layer system architecture for Unmanned Aerial Vehicle (UAV)-assisted offloading: Game-theoretic allocation and CTDE-MADDPG control
Yonglin Zhou, Qiu Fang, Yaonan Wang 0001 |
J. Syst. Archit. | 4 |
| 2026 | ZUMA: Training-Free Zero-Shot Unified Multimodal Anomaly DetectionabstractMultimodal anomaly detection (MAD) aims to exploit both texture and spatial attributes to identify deviations from normal patterns in complex scenarios. However, zero-shot (ZS) settings arising from privacy concerns or confidentiality constraints present significant challenges to existing MAD methods. To address this issue, we introduce ZUMA, a training-free, Zero-shot Unified Multimodal Anomaly detection framework that unleashes CLIP's cross-modal potential to perform ZS MAD. To mitigate the domain gap between CLIP's pretraining space and point clouds, we propose cross-domain calibration (CDC), which efficiently bridges the manifold misalignment through source-domain semantic transfer and establishes a hybrid semantic space, enabling a joint embedding of 2D and 3D representations. Subsequently, ZUMA performs dynamic semantic interaction (DSI) to enable structural decoupling of anomaly regions in the high-dimensional embedding space constructed by CDC, where natural languages serve as semantic anchors to help DSI establish discriminative hyperplanes within hybrid modality representations. Within this framework, ZUMA enables plug-and-play detection of 2D, 3D or multimodal anomalies, without training or fine-tuning even for cross-dataset or incomplete-modality scenarios. Additionally, to further investigate the potential of the training-free ZUMA within the training-based paradigm, we develop ZUMA-FT, a fine-tuned variant that achieves notable improvements with minimal parameter trade-off. Extensive experiments are conducted on two MAD benchmarks, MVTec 3D-AD and Eyecandies. Notably, the training-free ZUMA achieves state-of-the-art (SOTA) performance on both datasets, outperforming existing ZS MAD methods, including training-based approaches. Moreover, ZUMA-FT further extends the performance boundary of ZUMA with only 6.75 M learnable parameters. Yunfeng Ma, Min Liu 0008, Jingyu Zhou, Yuan Bian 0002, Yaonan Wang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2026 | Diffusion-Driven Self-Supervised Learning for Shape Reconstruction and Pose EstimationabstractFully-supervised category-level pose estimation aims to determine the 6-DoF poses of unseen instances from known categories, requiring expensive manual labeling costs. Recently, various self-supervised category-level pose estimation methods have been proposed to reduce the requirement of the annotated datasets. However, most methods rely on synthetic data or 3D CAD model, and they are typically limited to addressing single-object pose problems without considering multi-objective tasks or shape reconstruction. To overcome these challenges and limitations, we introduce a diffusion-driven self-supervised network for multi-object shape reconstruction and categorical pose estimation, only leveraging the shape priors. Specifically, to capture the SE(3)-equivariant pose features and 3D scale-invariant shape information, we present a Prior-Aware Pyramid 3D Point Transformer. This module adopts a point convolutional layer with radial-kernels for pose-aware learning and a 3D scale-invariant graph convolution layer for object-level shape representation. Furthermore, we introduce a Pretrain-to-Refine Self-Supervised Training Paradigm to train our network. It enables proposed network to capture the associations between shape priors and observations, addressing the challenge of intra-class shape variations by utilising the diffusion mechanism. Extensive experiments conducted on four public datasets and a self-built dataset demonstrate that our method significantly outperforms state-of-the-art self-supervised category-level baselines and even surpasses some fully-supervised instance-level and category-level methods. The project page is released at Self-SRPE. Yaonan Wang 0001, Mingtao Feng, Chao Ding 0006, Zheng Shou 0001, Ajmal Mian |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2026 | Distilling Object Detectors via Monte Carlo DropoutabstractKnowledge distillation (KD) has become a fundamental technique for model compression in object detection tasks. The data noise and training randomness may cause the knowledge of the teacher model to be unreliable, referred to as knowledge uncertainty. Existing methods neglect this uncertainty, potentially hindering the student's capacity to capture and understand latent "dark knowledge". In this work, we introduce a novel strategy that explicitly incorporates knowledge uncertainty, named Uncertainty-Driven Knowledge Extraction and Transfer (UET). Given the unknown, high-dimensional nature of the knowledge distribution, we employ Monte Carlo dropout to effectively estimate the teacher's uncertainty. Leveraging information theory, we combine uncertainty with deterministic knowledge, enabling the student to benefit from both precision and diversity. UET is a plug-and-play method that integrates seamlessly with existing distillation techniques. We validate our approach through comprehensive experiments across various distillation strategies, detectors, and backbones. Specifically, UET achieves state-of-the-art results, with a ResNet50-based GFL detector obtaining 44.1% mAP on the COCO dataset-surpassing baseline performance by 3.9%. Junfei Yi, Hui Zhang 0023, Jianxu Mao, Tengfei Liu 0005, Mingjie Li 0006, Sihao Lin, Hanyu Gu, Zhihui Li 0001, Xiaojun Chang, Yaonan Wang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 10 |
| 2026 | Efficient Point Cloud Processing With High-Dimensional Positional Encoding and Non-Local MLPsabstractMulti-Layer Perceptron (MLP) models are the foundation of contemporary point cloud processing. However, their complex network architectures obscure the source of their strength and limit the application of these models. In this article, we develop a two-stage abstraction and refinement (ABS-REF) view for modular feature extraction in point cloud processing. This view elucidates that whereas the early models focused on ABS stages, the more recent techniques devise sophisticated REF stages to attain performance advantages. Then, we propose a High-dimensional Positional Encoding (HPE) module to explicitly utilize intrinsic positional information, extending the "positional encoding" concept from Transformer literature. HPE can be readily deployed in MLP-based architectures and is compatible with transformer-based methods. Within our ABS-REF view, we rethink local aggregation in MLP-based methods and propose replacing time-consuming local MLP operations, which are used to capture local relationships among neighbors. Instead, we use non-local MLPs for efficient non-local information updates, combined with the proposed HPE for effective local information representation. We leverage our modules to develop HPENets, a suite of MLP networks that follow the ABS-REF paradigm, incorporating a scalable HPE-based REF stage. Extensive experiments on seven public datasets across four different tasks show that HPENets deliver a strong balance between efficiency and effectiveness. Notably, HPENet surpasses PointNeXt, a strong MLP-based counterpart, by 1.1% mAcc, 4.0% mIoU, 1.8% mIoU and 0.2% Cls. mIoU, with only 50.0%, 21.5%, 23.1%, 44.4% of FLOPs on ScanObjectNN, S3DIS, ScanNet, and ShapeNetPart, respectively. Yanmei Zou, Hongshan Yu, Yaonan Wang 0001, Zhengeng Yang, Xieyuanli Chen, Kailun Yang 0001, Naveed Akhtar |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2026 | A multimodal fusion model based on graph convolutional network for 3D neural morphology optimization
Hongji Qiu, Zhao Yao, Yaonan Wang 0001, Min Liu 0008 |
Pattern Recognit. | 4 |
| 2026 | FPS: Frequency prompt synchronization for micro-expression recognition
Jiateng Liu, Hengcan Shi, Yaonan Wang 0001, Wenming Zheng |
Pattern Recognit. | 3 |
| 2026 | Learning coherent matrixized representation in latent space for volumetric 4D generation
Qitong Yang, Mingtao Feng, Shijie Sun 0001, Weisheng Dong, Yaonan Wang 0001, Mian M. Ajmal |
Pattern Recognit. | 6 |
| 2026 | FaceEditor: Text-driven and mask-constrained face attribute editing
Lin Zhang 0041, Weiliang Meng, Paul L. Rosin, Yukun Lai, Yaonan Wang 0001 |
Pattern Recognit. | 7 |
| 2026 | CLEAR-MP: Clearance Learning-Based Efficient Motion Planning for Dual-Arm Robots Under End-Effector Orientation ConstraintsabstractDual-arm robotic manipulation of liquid or biochemical reagents poses critical challenges due to high-dimensional configuration spaces, stringent task-specific end-effector orientation requirements to prevent spillage, frequent inter-arm collisions, and cluttered experimental environments. This paper introduces CLEAR-MP (Clearance Learning-Based Efficient Motion Planning for Dual-Arm Robots under End-Effector Orientation Constraints), a modular framework that integrates multiple innovations: a decoupled learning-driven collision estimation module–comprising aPairwise Link Clearance Networkfor self-collision and aClearance Inference Networkfor environmental obstacles, aLearning-Driven Bidirectional Parallel Search Strategyfor accelerated tree expansion, parallel Cartesian batch sampling for efficient candidate generation, fast inverse-kinematics mapping, andLearning-Guided Batch Shortcut Optimizationto refine trajectories. Together, these components generate smooth, safety-certified paths with substantially reduced planning time and path length. Extensive simulations and real-robot experiments show that CLEAR-MP achieves an average path length of 2.391 m, average planning time of 3.529 s, outperforming state-of-the-art baselines by over 50% in computation and 40% in trajectory quality while maintaining strong generalization without retraining. Bo Chen 0047, Hui Zhang 0023, Yexin Fan, Yiming Jiang 0001, Chenguang Yang 0001, Yaonan Wang 0001 |
IEEE Trans Autom. Sci. Eng. | 7 |
| 2026 | Accelerating Outlier-Robust Point Cloud Registration by Known Gravity Directionsabstract3D point cloud registration, which seeks the optimal rigid transformation to align two point clouds, is a fundamental task in autonomous systems. However, the 3D correspondences between point clouds are prone to substantial outliers (mismatches), leading to significant decreases in registration accuracy. Existing outlier-robust registration methods commonly have high computational complexity and, hence, are limited in time-sensitive applications. Inertial measurement unit (IMU) sensors are widespread in modern autonomous systems and can offer precise gravity directions. Accordingly, we propose a highly efficient voting-based outlier removal method by leveraging the gravity prior in this paper. This pre-processing step can significantly reduce the candidate correspondence set for subsequent estimation, thus accelerating robust point cloud registration. We then leverage pairwise invariant features to decompose the optimization of rotation and translation. Further, we propose a two-stage consensus maximization solver to optimize the rotation and translation sequentially, leading to deterministic and robust registration. Extensive experiments on both synthetic and real-world datasets indicate that our method effectively boosts registration efficiency while exhibiting comparable robustness to state-of-the-art methods. Yinlong Liu, Yaonan Wang 0001 |
IEEE Trans Autom. Sci. Eng. | 4 |
| 2026 | HECE-IC: An Integrated Calibration Method for Delta Robot-Based Kitchen Waste Sorting SystemsabstractThis paper presents a multifunctional automated sorting system for kitchen waste based on a Delta robot. The sorting system is divided into three main modules: visual detection, information processing, and multifunctional robotic sorting. The visual detection module captures images of waste on a conveyor belt and transmits them in real-time to the information processing unit, where detection algorithms generate data on waste categories and grasp positions. The robotic arm, equipped with a force sensor, gripper, and suction cup, selects appropriate grasping or suction functions based on the waste type and shape to complete sorting. Additionally, this paper introduces a robust and efficient integrated calibration method for robot hand-eye-conveyor belt-encoder systems (HECE-IC), which enables simultaneous hand-eye and encoder calibration with only three simple steps. In simulations, the proposed method maintains a reconstruction error as low as 0.423 mm even under operation errors up to 0.6 mm. In practical experiments on the sorting platform, the average calibration error stabilized around 0.5 mm, achieving high calibration precision. The system achieved a sorting success rate of 90.2% and a sorting speed of 979 objects per hour. Our code is available at: https://github.com/TDA-2030/XRobot. Hai Qin, Songyun Deng, Qiaokang Liang, Dan Zhang 0006, Yaonan Wang 0001 |
IEEE Trans Autom. Sci. Eng. | 7 |
| 2026 | A Novel Neural-Network-Based MPC Framework for Whole-Body Motion Optimization and Control of Redundant Humanoid RobotsabstractThe rapid advancement of humanoid robotics has highlighted the challenge of coordinating high-redundancy whole-body motion and posture control under multiple safety constraints. The difficulty lies particularly in dual-arm task planning and the design of high-precision, real-time control strategies. To address these issues, this paper proposes a novel neural-network (NN) based real-time model predictive control (MPC) framework for humanoid robots. The framework innovatively integrates discrete recurrent neural networks (DRNN) with MPC, thereby extending their combined advantages to high-redundancy humanoid motion control. In addition, a primal-dual neural network (PDNN) solver is employed to compute multi-constrained kinematic MPC in a single iteration, avoiding the repeated iterations required in conventional MPC and significantly enhancing real-time performance. The proposed approach is rigorously validated through theoretical derivation and extensive experiments, including both numerical simulations and real-world trials. After theoretical verification in simulation, the NN-based MPC framework is deployed on a self-developed humanoid robotic platform. Experimental results confirm that the NN-based MPC method achieves effective and highly accurate whole-body task planning and real-time control, demonstrating its potential as a reliable solution for advanced humanoid robot control. Jie Wang 0091, Yiming Jiang 0001, Hui Zhang 0023, Yaonan Wang 0001 |
IEEE Trans Autom. Sci. Eng. | 5 |
| 2026 | SGTP-Net: Semantic Guidance and Texture Priors-Based Dual-Branch Segmentation Network for Surface Defect DetectionabstractDeep learning-based surface defect segmentation approaches have shown promising performance in recent years. However, segmenting defects with complex shapes, large variations in size, and weakly textured defects with indistinct characteristics still poses significant challenges. In this article, a novel semantic guidance and texture priors based dual-branch surface defect segmentation network (SGTP-Net) is proposed for those issues. Firstly, we construct a feature extraction network combines semantic and texture branches. The semantic branch establishes global contextual relationships, while the texture branch captures local features of defects, this dual-branch ensured the network to extract features from various complex defects. Secondly, we design a feature fusion strategy based on semantic guidance and texture priors. The semantic information is used to guides the output of texture branch. After that, the guided texture information provides valuable edge texture priors for each layers output in semantic branch. The two branches mutually guide each other for improving ability of weak textures feature extraction. Finally, we run our method on the NEU-Seg, MT-Defect and MSD datasets to conduct a comprehensive comparison with some state-of-the-art general object segmentation models and specialized surface defect segmentation methods. The experimental results show that our SGTP-Net performs well in surface defect detection, offering excellent semantic segmentation accuracy and exhibiting good stability and robustness in detecting various surface defects. Leqi Jiang, Liyue Ge, Chengzhong Wu, Yaonan Wang 0001, Ke Lu 0002, Congxuan Zhang |
IEEE Trans Autom. Sci. Eng. | 5 |
| 2026 | Event-Triggered Data-Driven Trajectory Tracking Control for Networked Mobile-Robot Systems: Application to Workpiece Transport
Xueming Zhang, Haoran Tan, Yaonan Wang 0001, Xin Wang 0003, Hui Zhang 0023, Zhongsen Wang, Jian Sun 0003 |
IEEE Trans Autom. Sci. Eng. | 3 |
| 2026 | Encoder-Only Image RegistrationabstractLearning-based techniques have significantly improved the accuracy and speed of deformable image registration. However, challenges such as reducing computational complexity and handling large deformations persist. To address these challenges, we analyze how convolutional neural networks (ConvNets) influence registration performance using the Horn-Schunck optical flow equation. Supported by prior studies and our empirical experiments, we observe that ConvNets play two key roles in registration: linearizing local intensities and harmonizing global contrast variations. Guided by these insights, we propose the Encoder-Only Image Registration (EOIR) framework comprising five modifications to existing approaches, to achieve a better accuracy-efficiency trade-off. EOIR separates feature learning from flow estimation, employing only a 3-layer ConvNet for feature extraction and a set of 3-layer flow estimators to construct a Laplacian feature pyramid, progressively composing diffeomorphic deformations under a large-deformation model. Results on six datasets across different modalities and anatomical regions demonstrate EOIR’s effectiveness, achieving superior accuracy-efficiency and accuracy-smoothness trade-offs. With comparable accuracy, EOIR provides better efficiency and smoothness, and vice versa. The source code of EOIR is available on Github. Xiang Chen 0008, Renjiu Hu, Min Liu 0008, Yaonan Wang 0001, Hang Zhang 0010 |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2026 | Fusing Events and Frames for Robust Gait RecognitionabstractReliable gait recognition under low-light conditions remains challenging for traditional cameras. Event cameras, with their high dynamic range and fine temporal resolution, offer a promising alternative but suffer from sparse signals under small motion amplitudes and vulnerability to noise or irrelevant movements (e.g., shadows). To address these issues, we propose FusionGait, a complementary fusion framework that integrates event with standard frames to achieve robust gait recognition. Specifically, we propose a Self-supervised Hierarchical Feature Extractor (SSL-HFE) built upon DINOv2, which employs learnable prompts to bridge the gap between gray frames and RGB frames, extracts multi-level semantic features, and enhances their discriminability through a self-supervised learning strategy. Then, we introduce the Complementary Fusion Learning Module (CFLM), which employs cross-cost volumes to explicitly model pixel-level correlations between frames and events, enabling effective cross-modal interaction and fusion. Furthermore, we propose EvGSimulator, a sensor-specific data augmentation strategy that simulates diverse illumination conditions based on physical properties. The framework maintains robustness even when real frames are unavailable by reconstructing frame-like representations from events, and it scales to ultra–high-frame-rate scenarios with hundreds of frames per second. We also collect DAVIS346-Gait-RGE, the multi-view semi-indoor & outdoor gait dataset captured with a DAVIS346 event camera, including three modalities: event streams, gray frames, and reconstructed frames. Experiments across multiple datasets show FusionGait achieves state-of-the-art performance, effectively surpassing single-modality methods. Liaogehao Chen, Zhenjun Zhang, Changchang Li, Yaonan Wang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2026 | Exploring Volume Representation Similarity in Long-Tail Biased Stereo Matching
Renjie Ding, Yaonan Wang 0001, Min Liu 0008, Jiazheng Wang 0001, Wenting Shen, Zhe Zhang 0022, Xiang Chen 0008 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2026 | Beijing Institute of TechnologyCMANet: A TCN-RMamba-Attention Network for Surgical Phase Online Recognition
Wenpei Fan, Yaonan Wang 0001, Licheng Liu, Min Liu 0008 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2026 | Spatial Multimodal Knowledge-Driven 3D Scene Graph Prediction With Vision-Language ModelabstractIn-depth understanding of 3D environments not only involves locating and recognizing individual objects but also requires inferring the relationships and interactions among them. However, most existing methods heavily rely on scene-specific contents, which leads to poor performance due to the noisy, cluttered, and partial nature of real-world 3D scenes. In this work, we find that the inherently hierarchical structures of 3D environments, derived from support relationships, aid in the automatic association of semantic and spatial arrangements of objects and provide rich geometric and topological information independent of specific scenarios. To this end, we propose a 3D scene graph generation model that leverages the hierarchical structures of 3D environments as spatial multimodal knowledge to enhance 3D scene graph generation. Specifically, we first devise a cross-modal tuning approach, where a visually-prompted vision language model is learned to infer the support relationships between objects in a low-resource way. Subsequently, we build a hierarchical visual graph and hierarchical symbolic knowledge graph using the fine-tuned vision language model to extract contextualized visual contents and relevant textual facts, respectively. Finally, we progressively accumulate 3D spatial multimodal knowledge about the hierarchical structures by correlating contextualized visual contents and textual facts using a novel graph reasoning network. In addition, to better evaluate the performance of 3D scene graph generation models, we propose a new benchmark 3DSSG-M by reorganizing the widely-used 3D scene graph generation dataset 3DSSG. This reorganization balances the predicate distribution of 3DSSG and reduces the influence of frequency bias. Extensive results and ablations attest to the effectiveness of the hierarchical structures in 3D environments and demonstrate the superiority of our proposed method over current state-of-the-art competitors. Haoran Hou, Mingtao Feng, Yulan Guo, Yaonan Wang 0001, Ajmal Mian |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2026 | DDIP: Mutual-Regularized Dual Deep Image Prior for Self-Supervised Compressive Spectral ImagingabstractIn recent years, numerous hyperspectral image (HSI) reconstruction methods have been proposed to enhance the imaging quality of coded aperture snapshot spectral compressive imaging (CASSI) systems. Among these methods, self-supervised Deep Image Prior (DIP)-based approaches have gained attention for their ability to reconstruct three-dimensional (3D) HSIs without the need for external training data. However, DIP methods often suffer from overfitting to high-frequency noise during the optimization process, leading to artifacts and loss of fine details. To address these challenges, we propose a Mutual-Regularized Dual Deep Image Prior (DDIP) framework that employs implicit mutual regularization between two DIP networks. By encouraging mutual constraints, DDIP effectively mitigates high-frequency learning bias and suppresses noise amplification. Additionally, we employ a Half Quadratic Splitting (HQS) optimization strategy to ensure stable and efficient convergence, progressively integrating complementary information from the dual networks. We provide a comprehensive convergence analysis of the DDIP framework and establish theoretical conditions to guide the progressive fusion of the dual networks, ensuring robust and reliable reconstruction. Based on the insights from the convergence analysis, we introduce an Adaptive Deep Image Prior inner-loop strategy that dynamically adjusts the inner-loop updates, ensuring balanced learning of low- and high-frequency components. Moreover, a Residual Spectral-Spatial Feature Attention Network (SSFAN) is designed to enhance spectral-spatial feature extraction, further improving reconstruction accuracy. Extensive experiments on benchmark datasets demonstrate that DDIP achieves competitive HSI reconstruction quality compared to state-of-the-art unsupervised and self-supervised methods. Lizhu Liu, Yaonan Wang 0001, Yurong Chen 0003, Hui Zhang 0023 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2026 | Uncertainty-Adaptive Volume for Unsupervised Homography EstimationabstractEstimating homography from an image pair is crucial for image alignment, and unsupervised methods that optimize feature reprojection error between target and warped source images have gained attention for their promising performance. In real-world scenes with multiple planes, such as moving objects, outlier rejection strategies are essential to mitigate the influence of non-dominant planes. Existing methods address this by learning a mask based on reprojection error, where high errors indicate non-dominant planes misaligned by homography. However, this error-fitting mask often overextends to the dominant plane, limiting the use of valid image regions for accurate estimation. This paper proposes a novel unsupervised method to compactly exclude non-dominant planes by introducing an uncertainty-adaptive cost volume for homography estimation. We first model uncertainty by assuming image features follow a Gaussian distribution derived from a prior Normal Inverse-Gamma distribution. The network-learned distribution parameters disentangle aleatoric uncertainty, distinguishing data-dependent errors within the total reprojection error. This uncertainty reflects inherent observation noise in image data, effectively indicating non-dominant planes. We then integrate this aleatoric uncertainty into the concatenation volume across image feature maps, creating an adaptive volume that filters out unreliable matching costs associated with non-dominant planes. This adaptive volume simplifies learning homography from the rich, redundant content in the concatenation volume, enabling more efficient and accurate estimation. Experiments demonstrate that our method outperforms existing approaches, achieving state-of-the-art performance both qualitatively and quantitatively. Jianqiao Luo, Yaonan Wang 0001, Mingtao Feng, Zhen Zhou 0003, Xuebing Liu, Yang Mo |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2026 | UniSurg: A Unified Multitask Framework for Robotic Surgical Scene UnderstandingabstractSurgical scene understanding is a vital intelligent technique in robot-assisted surgery, including surgical instrument detection, segmentation, and instrument–tissue interaction detection. Existing methods typically address these tasks in isolation, neglecting the intrinsic correlations among them. In this work, we innovatively propose a unified multitask framework named UniSurg, being the first to jointly address these three critical aspects of surgical scene understanding, thereby providing the robot with multidimensional perceptual capabilities. By exploring the inter-task correlations and reusing shared features, UniSurg has been demonstrated to significantly enhance the scene analysis performance. To address pose variability of the instruments under the constrained field of view in laparoscopic surgery, we design an Attention Enhanced Conditional Convolution (AEC-Conv) that dynamically adjusts kernels based on pose-specific features for improved adaptability. To further enhance interaction detection, we propose the Temporal Difference Enhancement module (TDE), which captures motion cues by amplifying inter-frame differences, and the Pyramid Global Feature Enhancement module (PGFE), which leverages graph-based hierarchical context to model global relational dependencies. Experiments on the Endovis2018 dataset and a clinical multitask dataset MILVis demonstrate the superior multitask performance of UniSurg. Wenting Shen, Yaonan Wang 0001, Min Liu 0008, Jiazheng Wang 0001, Renjie Ding |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2026 | Open-IndDet: Advancing Open-Set Industrial Surface Defect Detection via Robust Class-Unique Feature Representation
Zhen Yang 0026, Tianyong Zheng, Xuefeng Ni, Zhi Yan 0002, Yaonan Wang 0001, Leyuan Fang |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2026 | High-Precision Multi-Instance Registration for Stacked Objects in Bin-Picking ScenesabstractIn industrial bin-picking, robotic systems must estimate the poses of multiple object instances, where accurate pose estimation is essential for reliable downstream manipulation and grasping. Most existing multi-instance registration methods primarily establish point correspondences based on local features to alleviate the challenges posed by occlusion and clutter. However, local features are easily disturbed by neighboring instances and lack global context, leading to unreliable correspondences and degraded registration accuracy. In addition, the absence of rotational invariance further reduces correspondence accuracy in scenes with stacked instances and highly varying object orientations. To address these challenges, we present a one-stage multi-instance point cloud registration framework for stacked-object scenes. Our framework incorporates a rotation-invariant operator to enhance the robustness of feature representations under arbitrary orientations. Then, we propose a Center-Aware Res-Masked Transformer module, which incorporates an object center embedding to enrich global instance-level context and a center-aware residual mask prediction module to balance weight distribution across objects of varying sizes during training. Extensive experiments on the challenging ROBI dataset demonstrate that our method outperforms the competitive baseline MIRETR by more than 10% in mean precision, highlighting its effectiveness in complex bin-picking scenes. Furthermore, evaluations on the unstacked Scan2CAD dataset confirm the generalizability of the proposed framework across different application scenarios. Jiawen Zhao, Qing Zhu 0003, Yaonan Wang 0001, Weixing Peng, Jianxu Mao, Min Liu 0008, Xuebing Liu, Hui Zhang 0023 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2026 | SCAP: Semantic Prototype Alignment for Robust Point Cloud RegistrationabstractPoint cloud rigid registration is a fundamental problem in robotics, 3D reconstruction, and augmented reality. However, existing methods predominantly rely on local geometric neighborhoods, which fail to capture higher-order semantic structures and thus degrade performance under noisy or complex geometry conditions. To address these limitations, we propose SCAP, a new point cloud registration paradigm that transforms feature interaction from geometry-driven to semantics–geometric co-driven. Specifically, a semantic prototype extractor is devised to abstract high-level semantic prototypes through graph embedding and clustering, thereby mitigating sensitivity to local feature noise. Since semantic abstraction alone cannot guarantee consistent correspondences across point clouds, SCAP performs a prototype alignment path learning to infer reliable semantic mappings through optimal transport. To enhance cross-layer feature integration and prevent redundant attention, an alignment-driven cross-layer transformer is proposed to incorporate the learned priors into the attention mechanism, thereby enabling feature aggregation with improved semantic coherence and local precision. Extensive experiments on ModelNet, ModelLoNet, 3DMatch, and 3DLoMatch demonstrate that our SCAP consistently surpasses state-of-the-art approaches, showing superior robustness and generalization in challenging scenarios with noise and partial overlap. The code will be available at https://github.com/Zhou-111jy/SCAP.git. Jingyu Zhou, Yunfeng Ma, Yaonan Wang 0001, Min Liu 0008 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2026 | Cross-View Dynamic Learning-Based Multi-Class Industrial Anomaly DetectionabstractIndustrial anomaly detection plays a crucial role in smart manufacturing. Traditional methods typically train separate models for each category, leading to substantial memory demands and computational cost. Moreover, relying solely on single-view images is prone to detection blind spots and poor sensitivity to subtle defects. To address these problems, this study proposes CVDL, a cross-view dynamic learning-based multi-class industrial anomaly detection method. Specifically, the CVDL leverages a proposed cross-view dynamic attention in conjunction with intra-view self-attention to dynamically modulate the model’s attention on multi-view information, thereby enhancing the detection performance of subtle defects. Furthermore, a category-guided prompt is developed to utilize object category information, which improves the model’s class-aware detection accuracy. To enhance the model’s robustness, we introduce a structured noise injection strategy and a region-wise mask into the CVDL, mitigating the “identity shortcut” that preserves anomalies during reconstruction. Extensive experiments on the authentic multi-view industrial datasets (Real-IAD) and well-known datasets (MVTec-AD and VisA) confirm the superior detection capability and robustness of the proposed CVDL, and the overall performance of CVDL is superior to all advanced approaches on Real-IAD, achieving SoTA performance of 90.1% image-level and 99.0% pixel-level AUROC. The code will be available at https://github.com/zfinn1/CVDL.git. Jingyu Zhou, Yunfeng Ma, Yaonan Wang 0001, Min Liu 0008 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2026 | MGLD-TLNet: Multigeometric and Long-Distance Representation Network for Transmission Line InspectionabstractEffective transmission line (TL) inspection in complex corridor environments is essential for ensuring reliable power delivery. This work presents a 3D-based perception method for this task. The proposed method is designed by considering two key characteristics of TL inspection. First, the point cloud data are sparse and class distributions are highly imbalanced, which weakens the signals from thin conductors and tower components. To address this issue, we model long-range spatial relations along the corridor to mitigate data sparsity and imbalance. Second, strong structural correlations exist between conductors and towers, which can be leveraged to improve perception performance. To exploit this property, we construct a unified 3-D representation that jointly models towers, conductors, and vegetation, while fusing Cartesian and polar geometries through geometry-aware alignment. Experiments on real-world corridor datasets demonstrate that the proposed method, termed multigeometric and long-distance TL perception Network (MGLD-TLNet), consistently improves stability and accuracy under conditions of sparsity, occlusion, and complex environmental interactions. Hui Zhang 0023, Kaining Zhang, Baheti Biekezat, Hang Zhong, Junfei Yi, Jianxu Mao, Yaonan Wang 0001 |
IEEE Trans. Cybern. | 8 |
| 2026 | Adaptive Neural Network-Based Fault Detection for Thermal Process of Battery CellsabstractThis article presents an adaptive neural network (AdNN)-based fault detection framework for the thermal processes of lithium-ion (Li-ion) batteries governed by 2-D semilinear partial differential equations (PDEs) with partially-known dynamics. To address the challenges of unknown nonlinear heat generation and limited sensor measurements, a two-stage approach combining reduced-order modeling with adaptive neural observation is proposed. First, a computationally tractable reduced-order model is derived through spectral approximation techniques. An adaptive neural observer is then designed to simultaneously estimate battery states and unknown nonlinear dynamics using only available surface temperature measurements. For robust fault detection, a hybrid scheme is developed that integrates model-based residual generation with data-driven threshold generation. Experimental validation on a pouch-type battery demonstrates the effectiveness of the proposed method in reliably detecting thermal abnormalities. Yun Feng 0001, Ya-Zhi Zhang, Yaonan Wang 0001, Jun-Wei Wang 0001, Zhengguang Wu, Huaicheng Yan 0001, Han-Xiong Li |
IEEE Trans. Cybern. | 4 |
| 2026 | Contact Force Tracking Control for Aerial Manipulators in Unknown Dynamic EnvironmentsabstractIn this article, an admittance control strategy for aerial manipulators is proposed to achieve contact force tracking in unknown dynamic environments. First, considering the impact of variations in unknown environments on force tracking performance, an adaptive variable stiffness feature is incorporated into an advanced admittance model. The stiffness coefficient is dynamically adjusted using position and force feedback to generate the desired reference trajectory. Second, to address the issue of reference trajectory tracking under disturbances, a pose controller composed of a disturbance observer and barrier Lyapunov function is utilized to achieve stable tracking performance. In the absence of prior knowledge of disturbances, the state variables converge to a constrained range within a finite time, without introducing excessively high control gains. Finally, the stability of the proposed strategy is rigorously analyzed via Lyapunov tools. Both simulations and real-world experimental investigations are conducted to demonstrate the feasibility of the control strategy, highlighting its robust performance in maintaining a stable contact force during interaction with unknown dynamic environments. Zhiping Dai, Huimin Lu 0002, Hui Zhang 0023, Yaonan Wang 0001 |
IEEE Trans. Ind. Informatics | 7 |
| 2026 | GH-FMT: Heuristic-Based Expansion and Sampling for Fast Path Planning in Unknown EnvironmentsabstractMobile robots generally face formidable challenges in unknown dynamic environments, particularly in the aspects of rapidly searching feasible and high-quality solutions. In this article, Gaussian-heuristic fast marching tree (GH-FMT), including heuristic expansion and sampling, is proposed to provide high-quality solutions while simultaneously improving efficiency of planning and replanning in unknown dynamic environments. Specifically, an adaptive expansion region is defined using a boundary-based Gaussian distribution, minimizing potential redundant collision checking and supports effective and comprehensive exploration of the environments. A hybrid incremental search strategy is designed to reduce sample density and prioritized exploration of the search tree in promising directions using a heuristic single sampling method. Moreover, a hybrid optimization strategy is employed under time constraints to identify potential nodes, enabling efficient convergence toward high-quality solutions. Finally, through a series of challenging simulation scenarios and real-world experimental investigations, and by benchmarking against current-leading variants in the sampling-based planning class, the proposed GH-FMT demonstrated favorable advantages in flexibility, safety, and rapidity. It is shown that the search time is reduced on average by 15.8% and the path cost is reduced by on average 13.5%. Zhennan Lai, Zhaoguo Zeng, Wensheng Jiang, Huimin Lu 0002, Yaonan Wang 0001 |
IEEE Trans. Ind. Informatics | 7 |
| 2026 | Persistent Space-Level Path Segment Finding for Multiple Nonholonomic AgentsabstractMultiagent path finding (MAPF) is a critical problem in real-world multiagent systems, where agents reach their respective destinations without colliding with one another. Most MAPF solvers assume agents move at constant speeds with no delays and stop upon reaching their goal. However, real-world nonholonomic agents face delays like deceleration, lane changes, and acceleration, which impact efficiency. Moreover, tasks are continuously generated, making it challenging to meet production requirements. To address this issue, we propose a novel approach, the space-level path segment (SLPS) finding algorithm, which bridges the gap between traditional MAPF methods and the real-world requirements of nonholonomic agents. SLPS defines necessary and optional constraints to construct the collision detection graph (CDG) and formulates an optimization problem to minimize task completion time, resulting in the simplified CDG. It then computes a single SLPS, allowing nonholonomic agents to operate at varying speeds without synchronization. In addition, SLPS supports persistent MAPF tasks through four event-triggered methods, enabling flexible replanning in response to dynamic changes. Experimental results demonstrate that SLPS reduces average task completion time compared to the closest state-of-the-art solvers, proving its efficiency for real-world applications. Hongkai Fan, Bo Ouyang, Yaonan Wang 0001, Zhi Yan 0002, Qin Tan, Zhiheng Yao, Jiawen He |
IEEE Trans. Ind. Informatics | 3 |
| 2026 | PANDA: Progressive Adaptive Network for Defect-Aware Few-Shot SegmentationabstractFew-shot semantic segmentation aims to reduce reliance on dense annotations, while enhancing generalization to unseen categories. However, most methods are constrained by static global prototypes, which fail to represent subtle defect details, resulting in pronounced support–query misalignment. To address this issue, we propose progressive adaptive network for defect-aware few-shot segmentation (PANDA), a few-shot segmentation framework that integrates representational modulation with semantic consistency constraints. Specifically, to capture subtle and scale-sensitive variations in defect patterns, we design anchored representational modulation (ARM), which overcomes the rigidity of static prototypes by dynamically adjusting representations. In addition, we develop hierarchical semantic coherence (HSC), which enforces consistency across representation hierarchies to suppress the accumulation of semantic drift as depth increases. Collectively, ARM and HSC mitigate support–query misalignment and stabilize representations in few-shot defect segmentation. PANDA achieves state-of-the-art performance on MetFS-18, with 55.1% and 56.3% mean intersection over union under the one-shot and five-shot settings. Moreover, PANDA has been integrated into a real-time industrial inspection platform, where it delivers accurate segmentation across diverse defect types, highlighting its robustness in practical application. Yunfeng Ma, Min Liu 0008, Xiangfei Meng, Yaonan Wang 0001 |
IEEE Trans. Ind. Informatics | 7 |
| 2026 | Photovoltaic Module Inspection Based on Electromagnetic Induction-Assisted Scanning Photoluminescence Imaging
Yunze He, Baoyuan Deng, Cai Guo, Ruizhen Yang, Hong Zhang 0003, Yaonan Wang 0001 |
IEEE Trans. Ind. Informatics | 7 |
| 2026 | Unified Multimodal Industrial Anomaly Detection via Few Normal SamplesabstractMultimodal industrial anomaly detection (MIAD) is the process of integrating multiple sensor data and utilizing visual intelligence to identify abnormal states in industrial production. In this article, we focus on two main practical but challenging issues in MIAD, i.e., a unified model for multiclass anomaly detection, and model training with only few normal samples. The current mainstream “one-for-one” paradigm requires training time that grows exponentially, and it relies on a sufficient number of samples (even just normal samples), which cannot adapt to practical industrial scenarios with rich abnormal classes. To this end, we offer aUnifiedMIAD model that trained using onlyFew (e.g., 1, 2, and 4) normal samples, termed UniMF. Specifically, we propose a fusion-guided prompt engineering process that generates paired antithetical instance-specific prompts with the assistance of multimodal fusion at both query and token levels. To enable cross-modal prompt learning under multimodal conditions, UniMF performs multi-proxy pairwise matching that involves alignment among multimodal feature patches, embeddings, and tokens of antithetical prompts. Experimental results show that UniMF stands state-of-the-art performance while remaining “one-for-all” paradigm, and even outperforms “one-for-one” methods under certain settings. Cross-dataset evaluation between MVTec 3D-AD and Eyecandies datasets also shows the transferability of UniMF. Yunfeng Ma, Jingyu Zhou, Yaonan Wang 0001, Min Liu 0008 |
IEEE Trans. Ind. Informatics | 4 |
| 2026 | Micro Surface Defect Inspection of Aero-Engine Blades via Dynamic Cross-Scale Semantic Aggregation
Kaijie Li, Jingyu Zhou, Xiangfei Meng, Yaonan Wang 0001, Min Liu 0008 |
IEEE Trans. Ind. Informatics | 7 |
| 2026 | Investigating a Unified 3-D Object Detection Method for Different Multibeam LiDAR
Ziming Tao, Jianxu Mao, Yaonan Wang 0001, Caiping Liu, Junfei Yi, Zhenyu He 0015, Xiaojun Chang, Hui Zhang 0023 |
IEEE Trans. Ind. Informatics | 3 |
| 2026 | FPF: A Focused Perception Framework for Small Defect Identification in Complex Power Scenarios
Hui Zhang 0023, Baheti Biekezat, Yunkang Cao, Kaining Zhang, Tongzhi Niu, Yaonan Wang 0001 |
IEEE Trans. Ind. Informatics | 8 |
| 2026 | SAF: A Structure-Aware Framework for Radial Ice Thickness Detection on Overhead Transmission LinesabstractIce thickness estimation on overhead transmission lines (OHTL) is essential for mitigating icing-induced mechanical failures and ensuring safe grid operation. To address the challenges of detecting radial ice thickness in complex power line corridors, particularly geometric fragmentation of slender conductors and semantic ambiguity near occluded boundaries, this work proposes a structure-aware framework (SAF) based on 3-D point cloud segmentation and geometry-guided modeling. SAF introduces a structure-aware segmentation network, which integrates a cross-level spatial encoding module to preserve geometric continuity and a partition-aware loss to improve boundary localization under vegetation or tower occlusion. Building on accurate segmentation, a geometry-guided module performs centerline fitting and cross-sectional reconstruction to infer slice-level ice thickness. To support evaluation, a large-scale uncrewed aerial vehicle (UAV)-based point cloud dataset covering 32 OHTL is constructed, including six lines with ground-truth ice labels. Experimental results demonstrate that SAF achieves robust and accurate ice estimation across varied voltage levels and terrains, supporting its practical application in intelligent transmission line inspection and icing risk prevention. Hui Zhang 0023, Youyuan Tang, Yihong Cao, Kaining Zhang, Yunkang Cao, Tongzhi Niu, Jianxu Mao, Yaonan Wang 0001 |
IEEE Trans. Ind. Informatics | 9 |
| 2026 | Sparse-View 3-D Language Gaussian Splatting for Zero-Shot Robotic Graspingabstract3-D language Gaussian splatting has recently shown strong potential for open-vocabulary scene understanding and robotic manipulation. However, most existing methods require dense multiview observations to achieve accurate geometry reconstruction and reliable semantic alignment, which limits their applicability in scenarios where only sparse-view observations are available. In this work, we propose SparseGrasper, a framework for language-guided zero-shot robotic grasping under sparse-view conditions. SparseGrasper constructs a 3-D Gaussian language field from as few as three RGB images, enabling joint reasoning over geometry and semantics without the need for dense observations. To improve representation learning under sparse observations, we introduce a dual feature distillation module that fuses local object features with global contextual cues. We further design a language-guided grasp pose generation strategy that incorporates semantic grounding into grasp candidate selection, encouraging grasps that are both semantically relevant and geometrically feasible. Real-world experiments on a 7-DoF robotic manipulator validate that SparseGrasper effectively performs language-guided grasping of diverse, previously unseen objects from sparse observations. Yaonan Wang 0001, Wenrui Chen, He Xie, Zhengping Che, Pei Ren, Jian Tang 0008 |
IEEE Trans. Ind. Informatics | 2 |
| 2026 | Staged Modulation Diffusion Policy With Complementary Visual Fusion for Robotic Workpiece AssemblyabstractHigh-precision robotic assembly remains a challenge in intelligent manufacturing. Existing vision-based approaches often learn predefined trajectories from large image corpora yet underutilize task-relevant visual cues during execution, limiting deployment. We present a diffusion-based end-to-end assembly policy that performs redundancy-aware cross-scale fusion and provides stage-dependent conditioning for action generation. Specifically, we introduce a bidirectional-attention complementary visual fusion module that aligns cross-scale observations and produces scale-consistent features with reduced redundancy. We then introduce a dual-stream adaptive modulation module that enables time-state correlated routing and progressively shifts emphasis from scene-level stabilization to contact-level refinement within the denoising process. Coupling complementary visual fusion with adaptive-modulation-conditioned denoising, we develop a staged modulation diffusion policy for end-to-end action generation, producing temporally coherent and geometrically accurate actions. Experiments on five real-world tasks demonstrate consistent improvements over representative baselines, with average gains of 41.10% in overall task success rate and 63.16% in precise assembly rate on three representative assembly benchmarks. Yaonan Wang 0001, Mingtao Feng, Renjie Ding, Hui Zhang 0023 |
IEEE Trans. Ind. Informatics | 2 |
| 2026 | Second-Order Robust Iterative Pose Optimization for Fine-Grained Cross-View LocalizationabstractFine-grained cross-view localization seeks to estimate precise camera poses by matching ground images with GPS-tagged aerial imagery. Existing methods typically employ first-order iterative optimization to progressively update the camera pose based on cross-view feature correspondences. However, they rely on local features and neglect global and complementary contextual information, making them prone to local optima and slow convergence under large initial errors or strong disturbances. To overcome these limitations, we propose a second-order robust iterative pose estimation framework for fine-grained cross-view localization. Firstly, we devise a second-order deep iterative optimization module to capture complementary forward and backward motion cues, leading to a bidirectional correlation volume. A motion aggregator uses the volume to approximate the dynamics of second-order iterators, substantially facilitating convergence and robustness. In addition, a bidirectional motion-aware robust regularization module mitigates geometric distortions and outlier interference by leveraging bidirectional motion cues to generate fine-grained confidence maps, adaptively suppressing unreliable regions and enhancing the stability of iterative optimization and pose estimation accuracy. Extensive experiments demonstrate that the proposed framework achieves faster convergence and higher pose estimation accuracy than state-of-the-art methods, particularly under large initial errors and challenging conditions. Mingtao Feng, Jianqiao Luo, Weisheng Dong, Yaonan Wang 0001, Ajmal Mian |
IEEE Trans. Image Process. | 5 |
| 2026 | NRSeg: Noise-Resilient Learning for BEV Semantic Segmentation via Driving World ModelsabstractBirds' Eye View (BEV) semantic segmentation is an indispensable perception task in end-to-end autonomous driving systems. Unsupervised and semi-supervised learning for BEV tasks, as pivotal for real-world applications, underperform due to the homogeneous distribution of the labeled data. In this work, we explore the potential of synthetic data from driving world models to enhance the diversity of labeled data for robustifying BEV segmentation. Yet, our preliminary findings reveal that generation noise in synthetic data compromises efficient BEV model learning. To fully harness the potential of synthetic data from world models, this article proposes NRSeg, a noise-resilient learning framework for BEV semantic segmentation. Specifically, a Perspective-Geometry Consistency Metric (PGCM) is proposed to quantitatively evaluate the guidance capability of generated data for model learning. This metric originates from the alignment measure between the perspective road mask of generated data and the mask projected from the BEV labels. Moreover, a Bi-Distribution Parallel Prediction (BiDPP) is designed to enhance the inherent robustness of the model, where the learning process is constrained through parallel prediction of multinomial and Dirichlet distributions. The former efficiently predicts semantic probabilities, whereas the latter adopts evidential deep learning to realize uncertainty quantification. Furthermore, a Hierarchical Local Semantic Exclusion (HLSE) module is designed to address the non-mutual exclusivity inherent in BEV semantic segmentation tasks. The proposed framework is evaluated on BEV semantic segmentation using data generated by multiple world models, with comprehensive testing conducted on the public nuScenes dataset under unsupervised and semi-supervised settings. Experimental results demonstrate that NRSeg achieves state-of-the-art performance, yielding the highest improvements in mIoU of 13.8% and 11.4% in unsupervised and semi-supervised BEV segmentation tasks, respectively. The source code will be made publicly available at https://github.com/lynn-yu/NRSeg. Siyu Li 0002, Yihong Cao, Kailun Yang 0001, Zhiyong Li 0001, Yaonan Wang 0001 |
IEEE Trans. Image Process. | 6 |
| 2026 | Information-Bottleneck-Guided Hybrid Neural Architecture Search for Temporal Action Detection in Untrimmed VideosabstractTemporal Action Detection (TAD) in untrimmed videos requires effective spatial feature extraction for precise action classification and temporal feature modeling for accurate boundary localization. To achieve effective spatio-temporal feature integration, several works manually design rule-based (i.e., sequential or parallel) hybrid Mamba-Transformer networks for TAD. However, few studies explore diverse integration strategies and network topologies due to the inherent limitations of manual design. Therefore, we propose NAS-TAD, the first Neural Architecture Search framework for TAD, systematically exploring this untouched problem. Specifically, we develop a spatio-temporal NAS objective function based on information-bottleneck theory to quantify task-relevant spatio-temporal features, providing interpretable guidance for the network search and optimization process. Furthermore, we reformulate Transformer self-attention as a state-space model, thereby enabling seamless switching between Mamba and Transformer blocks in a unified weight-sharing search space. Consequently, comprehensive experiments on ActivityNet, THUMOS14, HACS and FineAction demonstrate the effectiveness of the searched hybrid architectures, providing new insights into temporal and spatial feature fusion for TAD. Code is available for reproduction at https://github.com/tyhnu/nastad.git. Mansen Chen, Lepeng Chen, Min Liu 0008, Yaonan Wang 0001 |
IEEE Trans. Image Process. | 6 |
| 2026 | Semantic-Aware Multimodal Collaborative Learning for Unsupervised Visible-Infrared Person Re-IdentificationabstractUnsupervised visible-infrared person re-identification (VI-ReID) is challenging due to the significant modality gap between visible and infrared images. Most existing methods rely on one-hot clustering pseudo-labels as supervision signals, which often fail to capture the full semantic relationships among samples and are highly susceptible to noise. To address these limitations, we propose a Semantic-aware Multimodal Collaborative Learning (SAMCL) framework for unsupervised VI-ReID. Specifically, a Modality-aware Semantic Fusion (MSF) module is designed to bridge the inter-modality gap by integrating complementary semantic details from both visible and infrared modalities, generating enriched cross-modal supervision signals, for cross-modal collaborative learning. Meanwhile, we present a Dynamic Contrastive Learning (DCL) module to refine intra-modality feature learning by dynamically aligning samples with their neighboring centroids in the feature space, improving clustering reliability and intra-modality feature discrimination. By combining the two modules, SAMCL harnesses multimodal collaboration, minimizes dependence on noisy pseudo-labels, and provides a robust approach to unsupervised VI-ReID. Extensive experiments demonstrate the superiority of our proposed method. For instance, on the SYSU-MM01 dataset, our model achieves a Rank-1 accuracy of 68.68% in the All Search setting, surpassing the state-of-the-art (SOTA) by 3.48%. On the RegDB dataset, it achieves a Rank-1 accuracy of 94.47% in the Visible-to-Infrared setting, outperforming the SOTA by 3.57%. On the LLCM dataset, it achieves a Rank-1 accuracy of 50.6% in the Visible-to-Infrared setting, outperforming the SOTA by 3.7%. The code is available at https://github.com/luoshixi123/SAMCL. Shixi Luo, Min Liu 0008, Gautam Srivastava 0001, Shuai Liu 0002, Yaonan Wang 0001 |
IEEE Trans. Image Process. | 6 |
| 2026 | P3C-DNet: Pseudo-Groundtruth Contrastive Learning With Color Calibration Dehazing NetworkabstractExisting dehazing methods primarily rely on synthetic hazy images for supervised learning. While effective on synthetic datasets, these methods often struggle to generalize to real-world hazy images, leading to issues such as color distortion and incomplete haze removal. Moreover, their limited adaptability to real-world datasets and inability to handle complex haze scenarios remain significant challenges. To address these limitations, we propose a novel unsupervised framework P3C-DNet (Pseudo-groundtruth Contrastive learning with Color Calibration Dehazing Network). Our P3C-DNet introduces a Pseudo-groundtruth image generation strategy through the Pseudo-groundtruth Contrastive Supervision (PCS) module, which overcomes the lack of real haze-free training data by generating high-quality Pseudo-groundtruth images. To further refine the dehazing process, we incorporate a codebook-based image coding and matching mechanism that aligns Pseudo-groundtruth images with hazy inputs, enhancing the accuracy and detail of the dehazed outputs. To address the prevalent issue of color distortion, especially in complex environments, our P3C-DNet integrates a Dynamic Color Restoration Block (DCRB) to ensure visual quality and color consistency in the dehazed results. Experimental evaluations demonstrate that our P3C-DNet achieves superior performance in haze removal, color fidelity, and detail preservation, significantly outperforming existing methods and setting a new benchmark for real-world dehazing tasks. Ze Ouyang, Weiliang Meng, Paul L. Rosin, Yukun Lai, Yaonan Wang 0001 |
IEEE Trans. Image Process. | 6 |
| 2026 | You Can Only Tune Normalization: A Simple and Effective Approach to Parameter-Efficient Fine-TuningabstractTo tackle the issue of excessive parameter volumes during fine-tuning of large-scale pre-trained models with full parameters, Parameter-Efficient Fine-Tuning (PEFT) methods have been introduced. The core concept involves freezing the backbone network of the model and updating only a small subset of parameters. This strategy not only decreases the number of parameters needed for training but also delivers performance comparable to Full-Tuning, even surpassing it on certain datasets. However, most popular PEFT methods introduce extra parameters or modules for fine-tuning, which come with inherent limitations. In response, we propose a straightforward and efficient PEFT method called You Can Only Tune Normalization (YONO). YONO focuses solely on tuning the normalization layer and the final classification layer of the model. This method avoids adding extra modules, making it easily applicable to any model without causing inference delays. We extensively tested YONO on 28 benchmark datasets, and the results indicate that it requires significantly fewer parameters compared to other advanced PEFT methods. Additionally, we validated YONO’s efficiency and generalizability across various vision models. Finally, we further explore the essence of PEFT methods, whether they learn new knowledge or expose the capabilities that a model has already learned. Our findings suggest that YONO is more sensitive to improvements in dataset quality, making it a promising candidate for future scaling to larger models. Lingyun Huang, Jianxu Mao, Junfei Yi, Ziming Tao, Ziyang Peng, Wei He 0001, Rui Liu 0028, Yaonan Wang 0001 |
ACM Trans. Intell. Syst. Technol. | 8 |
| 2026 | Self-Expert Imitation With Purifying Latent Feature for Generalization in Visual Reinforcement LearningabstractThe generalization ability of visual reinforcement learning, which allows the policy trained in the source domain to guide agents in similar unknown target environments, is one of the cores applied to visual navigation and autonomous driving. Recently, methods such as data augmentation techniques, self-supervised learning methods, and the generative adversarial network were employed to enhance the generalization capability of policy neural networks in visual reinforcement learning. However, current state-of-the-art methods, after utilizing domain-general latent features to train the RL policy, result in the loss of certain state-specific features, leading to diminished policy performance following generalization. To tackle these challenges, we designed a technical framework called self-expert imitation with purifying latent features, which enables the trained policy to effectively guide agents in scenarios similar to the training environment, without compromising the performance of the policy-guided agent in task completion. Additionally, a novel method was developed for separating domain-general and domain-specific latent vectors based on a variational autoencoder, enabling the domain-general component to exhibit strong and stable zero-shot generalization performance in unseen visually similar domains. Extensive experiments on the CarRacing game demonstrated that our approach achieves strong and stable generalization performance in unseen environments, without compromising the performance of the policy in guiding agents to complete tasks. Lin Chen 0034, Yang Mo, Yaonan Wang 0001, Zhiqiang Miao, Kai Zeng 0010, Mingtao Feng, Zhen Zhou 0003, Sifei Wang, Danwei Wang |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2026 | A Two-Stage Co-Evolutionary Algorithm Enhanced by Reinforcement Learning for Efficient Aircraft Test Task AllocationabstractEfficient long-term flight test task allocation is crucial for aircraft development, yet it poses significant challenges due to the complex, extended timelines and numerous interdependent tasks. Current methods often rely on expert judgment, leading to inefficiencies, delays, and suboptimal planning. This paper addresses these challenges by introducing a reinforcement learning-enhanced two-stage evolutionary algorithm. The algorithm incorporates a mixed-integer nonlinear programming model aimed at minimizing the maximum flight test duration while adhering to multiple constraints. Key innovations include a two-stage evolutionary framework with optimized encoding schemes and crossover operators, and dynamic adjustment of crossover and mutation probabilities using Q-learning. These features enable the algorithm to handle the complexity and unpredictability of long-term flight test planning. Extensive experiments validate the approach, demonstrating its ability to generate high-quality, adaptive flight test plans in real-world, complex scenarios. Qiu Fang, Huiyou Zhu, Yi Mi, Yaonan Wang 0001 |
IEEE Trans. Intell. Transp. Syst. | 6 |
| 2026 | A Reinforcement Learning-Based Decentralized Control Strategy for Eco-Safe Mixed Platooning With CAVs and HDVsabstractMixed platoons, consisting of connected autonomous vehicles (CAVs) and human-driven vehicles (HDVs), are expected to dominate future roadways. However, achieving substantial improvements in fuel economy and safety under time-varying, mixed traffic conditions in real-world scenarios remains challenging. To this aim, this paper proposes a deep reinforcement learning (DRL) based decentralized control strategy, which consists of two levels: 1) At the methodology level, the hierarchical platoon formation algorithm systematically organizes the mixed platoons into local platoons with uniform car-following patterns and sub-platoons within a generic HDV-CAVs structure, enabling adaptation to time-varying traffic volume and mixed traffic heterogeneity. 2) At the operation level, the HDV-CAVs unit is embedded in the learning environment to capture uncertain HDV behaviors through state and reward propagation channels. Multiple objectives are explicitly integrated into the reward function. A safety-supervised decentralized proximal policy optimization algorithm is developed to enhance safety and training efficiency. Extensive validations demonstrate the effectiveness of the proposed strategy in improving fuel economy and safety under different traffic demand levels and penetration rates. Xiangcheng Pan, Xiaofang Yuan, Zhigang Ling, Zhe Li 0050, Yaonan Wang 0001 |
IEEE Trans. Intell. Transp. Syst. | 6 |
| 2026 | Remaining Useful Life Prediction for Key Components of Transportation Vehicles: A Physics-Informed PerspectiveabstractIn transportation systems, accurately estimating the remaining useful life (RUL) of critical components, such as aircraft engines, Battery Management Systems (BMSs), is crucial for the safe and reliable operation and manufacturing of transportation vehicles. However, most existing research overlooks the underlying physical information, which is vital for more precise RUL prediction. To fill this gap, this paper proposes a physics-informed method for predicting the RUL of key components of transportation vehicles. By integrating the Mamba network with a multi-head attention mechanism, we capture and emphasize key features and trends in the equipment’s operational state, improving prediction accuracy. Additionally, we introduce a Physics-Informed Neural Network (PINN) framework to model the underlying physical relationships between RUL and sensor data, incorporating these relationships as a regularization term in the loss function to enhance predictive capability and interpretability. We conducted experimental validation using the C-MAPSS aircraft engine dataset (operation) and the transportation vehicle chip manufacturing dataset (manufacture). The results show that the proposed method significantly improves the accuracy of RUL prediction, providing strong support for the intelligent maintenance and reliability management of key components in transportation vehicles. Qing Zhu 0003, Yucong Shi, Yun Feng 0001, Ya-Zhi Zhang, Haoran Tan, Yaonan Wang 0001, Wanke Yu, Yongfu Li 0001 |
IEEE Trans. Intell. Transp. Syst. | 6 |
| 2026 | Neural Optimization for Image Registration via Joint Modeling of Global Affine and Local Deformation TransformationsabstractConventional registration approaches frequently underperform when applied to sparse feature alignment (e.g., retinal vessels and filamentous collagen fibers in second-harmonic generation (SHG) and bright-field (BF) images), as these tasks demand simultaneous handling of global affine registration and local deformation correction. End-to-end learning-based approaches struggle with minimal effective gradients from loss back-propagation of these sparse features, while descriptor matching methods, though helpful, lack fidelity loss and fail to adapt to local deformation. To address these issues, we propose Neural Affine Optimization (NeOn), which implicitly approximates discrete optimization using a few neural network layers, combined with a sampling-regression layer to handle affine transformations. NeOn allows iterative refinement with fidelity loss and provides a flexible transition between a purely affine configuration and a linear weighted blend of affine and deformation fields. NeOn's performance was validated on four public datasets. In multi-modal SHG-BF microscopy registration, NeOn achieved top rankings on the validation leaderboard for Task 3 of the Learn2Reg Challenge 2024. For retinal image registration, NeOn outperformed existing methods on both mono-modal and multi-modal datasets, reducing target registration error from 6.3 to 2.1 pixels in mono-modal and from 2.6 to 1.8 pixels in multi-modal registration. Furthermore, NeOn demonstrates strong generalization and can be effectively extended to 3D multi-modality image registration scenarios. Xiang Chen 0008, Renjiu Hu, Jiacheng Wang 0001, Min Liu 0008, Yaonan Wang 0001, Jiazheng Wang 0001, Rongguang Wang, Gaolei Li, Hang Zhang 0010 |
IEEE Trans. Medical Imaging | 5 |
| 2026 | CiSeg: Unsupervised Cross-Modality Adaptation for 3D Medical Image Segmentation via Causal InterventionabstractUnsupervised domain adaptation (UDA) addresses the domain shift problem by transferring knowledge from labeled source domain data (e.g. CT) to unlabeled target domain data (e.g. MRI). While state-of-the-art methods reduce domain gaps via image- or feature-level alignment, their reliance on spurious correlations in the training data often limits generalization across domains. To overcome this limitation, we propose the Causal Intervention Segmentation Network (CiSeg), a novel framework that first integrates causal inference into UDA. A Structural Causal Model (SCM) is first constructed for the source domain to disentangle causal variables from bias variables, alleviating the impact of spurious correlations. Based on this SCM, we introduce a Counterfactual Disentanglement (CD) module to decompose the source domain's latent features into distinct causal and bias components, effectively eliminating their mutual dependencies. To enhance cross-domain consistency, two auxiliary components are introduced: Prototype-guided Contrastive Learning (PCL) and Causal-bias Residual Alignment (CBRA). PCL aligns pixel-level representations with their corresponding semantic prototypes, promoting stronger intra-class consistency and clearer inter-class separability. CBRA employs adversarial learning to align causal and bias residual features across domains, further enhancing feature-level invariance. Extensive experiments on cardiac, abdominal multi-organ, and BraTS18 segmentation tasks demonstrate that CiSeg outperforms state-of-the-art methods, achieving superior segmentation performance and robust cross-domain generalization. Code and models are available at https://github.com/lvpeiqing/CiSeg. Peiqing Lv, Yaonan Wang 0001, Min Liu 0008, Zhe Zhang 0022, Yunfeng Ma, Licheng Liu, Erik Meijering |
IEEE Trans. Medical Imaging | 2 |
| 2026 | EPDiff: Erasure Perception Diffusion Model for Unsupervised Anomaly Detection in Preoperative Multimodal ImagesabstractUnsupervised anomaly detection (UAD) methods typically detect anomalies by learning and reconstructing the normative distribution. However, since anomalies constantly invade and affect their surroundings, sub-healthy areas in the junction present structural deformations that could be easily misidentified as anomalies, posing difficulties for UAD methods that solely learn the normative distribution. The use of multimodal images can facilitate to address the above challenges, as they can provide complementary information of anomalies. Therefore, this paper propose a novel method for UAD in preoperative multimodal images, called Erasure Perception Diffusion model (EPDiff). First, the Local Erasure Progressive Training (LEPT) framework is designed to better rebuild sub-healthy structures around anomalies through the diffusion model with a two-phase process. Initially, healthy images are used to capture deviation features labeled as potential anomalies. Then, these anomalies are locally erased in multimodal images to progressively learn sub-healthy structures, obtaining a more detailed reconstruction around anomalies. Second, the Global Structural Perception (GSP) module is developed in the diffusion model to realize global structural representation and correlation within images and between modalities through interactions of high-level semantic information. In addition, a training-free module, named Multimodal Attention Fusion (MAF) module, is presented for weighted fusion of anomaly maps between different modalities and obtaining binary anomaly outputs. Experimental results show that EPDiff improves the AUPRC and mDice scores by 2% and 3.9% on BraTS2021, and by 5.2% and 4.5% on Shifts over the state-of-the-art methods, which proves the applicability of EPDiff in diverse anomaly diagnosis. The code is available at https://github.com/wjiazheng/EPDiff. Jiazheng Wang 0001, Min Liu 0008, Wenting Shen, Renjie Ding, Yaonan Wang 0001, Erik Meijering |
IEEE Trans. Medical Imaging | 5 |
| 2026 | EdinoGait: Transferring Large Visual Models to Event-Based Vision for Enhancing Gait RecognitionabstractCurrent gait recognition methods heavily rely on various gait representations (e.g., silhouette sequences) generated by task-specific, supervised upstream processes, which inevitably incur high annotation costs and the risk of cumulative errors. Recently, generic knowledge from task–agnostic large visual models (LVMs) has been successfully applied to gait recognition, freeing the field from such dependencies. However, this approach does not address challenges posed by traditional cameras in handling scenarios with low latency, high speed, and high dynamic range. In this paper, we introduce EdinoGait, a novel and effective gait recognition framework that leverages event-based LVMs to overcome the scarcity of large-scale event-based datasets. Specifically, due to the distinct modality gap between image and event data and the lack of large-scale datasets, transferring LVMs to event-based vision is non-trivial. To address this, we introduce a novel event encoder that mitigates the modality gap through event prompts and a$CLS$patch contrastive loss. Subsequently, we design an autoencoder-based dual-alignment module to eliminate background noise brought by LVMs while preserving the motion details provided by event data. Additionally, to promote the application of event cameras in gait recognition, we collect the first semi-indoor, multi-view gait dataset captured by the DAVIS346 event camera. This dataset comprises 6,150 sequences (two modalities: grayscale images and event streams) of 41 subjects captured under two lighting conditions and five view angles ($0^{\circ }$,$45^{\circ }$,$90^{\circ }$,$135^{\circ }$, and$180^{\circ }$). Specifically, for each lighting condition and viewing angle, there are six sequences representing normal walking (NM), three representing walking with a backpack (BG), three with a portable bag (PT), and three with a coat (CL). Comprehensive experiments conducted on our event-based gait dataset and EV-CASIA-B demonstrate that EdinoGait significantly outperforms frame-based LVMs. Notably, under low-light conditions, the recognition accuracy of frame-based LVMs declines sharply, while EdinoGait exhibits robust performance. Liaogehao Chen, Zhenjun Zhang, Yaonan Wang 0001 |
IEEE Trans. Multim. | 3 |
| 2026 | BCNet: Butterfly-Shaped Convolutions Network for Lightweight Edge DetectionabstractAiming at the multi-task optimization conflicts (structure-detail-denoising coupling) and high computational costs caused by existing edge detectors' reliance on complex pre-trained models, this paper proposes BCNet - the first innovative framework that synergistically integrates biological visual mechanisms and information theory to achieve triple decoupling. First, inspired by butterfly-shaped receptive fields in the visual system, we design learnable butterfly-shaped convolution kernels as fundamental operators. These kernels inherently enhance structural perception and suppress noise without requiring deep architectures and pre-training, achieving structure-denoising separation. Second, to address the detail-noise coupling issue in existing methods that reconstruct edge images from multi-scale downsampled features, we propose a conditional entropy-based uncertainty modeling approach. The uncertainty of feature loss during downsampling is quantified via Gaussian distributions, while a self-supervised mechanism dynamically assigns detaillearning weights, enabling the model to learn more details from high-uncertainty regions while avoiding noise introduction. With at most about 2M parameters and real-time inference speeds up to 152 FPS, BCNet demonstrates strong competitiveness across four benchmark datasets, providing a novel, efficient, and lightweight solution for edge detection.https://github.com/StarkLuo/BCNetCode will be available. Zhengqiao Luo, Zhenjun Zhang, Chuan Lin 0003, Yaonan Wang 0001 |
IEEE Trans. Multim. | 4 |
| 2026 | DPAKS: Reliable DETR Guided by Prior Auxiliary Knowledge for Small Object DetectionabstractDeep neural networks are highly effective at transforming sparse and unstructured data into dense and semantic representations, demonstrating strong capabilities in object detection tasks. However, their performance often diminishes when detecting small-sized objects due to the loss or corruption of critical information during feature extraction. To address this challenge, DPAKS is introduced, a reliable DETR- like detector for small objects enhanced with directional prior auxiliary knowledge to guide the model's focus on small objects. In the decoder of DPAKS, a small denoising training strategy is employed that reduces the interference of noisy queries generated from real small objects. This approach effectively learns the features of small objects during the denoising process, sharpening the model's attention to small-sized objects. Additionally, to enhance the reliability of DPAKS's backbone, an auxiliary branch is introduced that provides supervision via shorter paths, improving the optimization of low-level feature parameters. This branch facilitates the transmission of gradient information suited for small objects without interfering with the detection of other sized objects. Furthermore, a new supervision head is proposed and added to the detection head of DPAKS, which categorizes object sizes based on artificial prior knowledge. This guides the model to effectively learn size categories and become more sensitive to small objects. Remarkably, DPAKS achieves competitive performance in small object detection without imposing additional computational burdens at the inference stage. The code is available athttps://github.com/XUhaozhi88/DPAKS. Haozhi Xu, Xiaofang Yuan, Yaonan Wang 0001 |
IEEE Trans. Multim. | 3 |
| 2026 | Open Set Industrial Surface Defect Recognition With High Frequency Feature Enhancement and Class Mutual-Information ConstraintabstractDefect detection in multimedia data plays a pivotal role in industrial manufacturing. However, existing methods are primarily designed for closed-world scenarios and can only identify defect classes in the training data, limiting their ability to effectively detect unknown class defects that arise during production. To address this critical limitation, we propose a novel approach by introducing industrial defect open set recognition (IDOSR), which overcomes the challenge of recognizing unknown defect classes. Furthermore, to tackle the issues of limited training samples and subtle inter-class differences in IDOSR, we present a high-frequency feature enhancement open set recognition (HFFE-OSR) method. Specifically, HFFE-OSR employs a high-frequency structural feature fusion enhancement strategy to meticulously extract and fuse defect-related high-frequency structural features. This enables the network to comprehensively learn defect target representations even under limited training samples, resulting in robust feature extraction for known classes, thereby improving the discriminability between known and unknown classes and addressing the difficulty of distinguishing between them. Additionally, a class mutual information constraint strategy is introduced to measure and reduce the mutual information among defect features from different classes. This ensures the independence of defect features across known classes, further enhancing their discriminability and significantly improving recognition performance for known classes. Extensive experiments demonstrate that the proposed method significantly outperforms state-of-the-art OSR methods on ID-OSD and MVTec datasets, achieving improvements of at least 7% in accuracy (ACC), 22% in F1 score, and 10% in AUROC, highlighting the effectiveness of our approach in industrial defect detection. Zhen Yang 0026, Tianyong Zheng, Xuefeng Ni, Zhi Yan 0002, Shangzhi Liu, Yingtian Yu, Yaonan Wang 0001, Leyuan Fang |
IEEE Trans. Multim. | 8 |
| 2026 | Image-Quality-Guided Consistency-Alignment Network for Fault DiagnosisabstractSignal-to-image transformation has been widely used in mechanical fault diagnosis to provide unified visual representations for downstream diagnostic models. However, the resulting diagnostic images can exhibit substantial quality variations caused by measurement noise, sensor degradation, and partial signal loss. Most existing methods either discard low-quality samples or implicitly assume equal reliability across samples, leading to information loss or degraded robustness. This paper proposes an image-quality-guided consistency-alignment network that explicitly estimates sample quality and leverages degraded data during training. First, vibration signals are decomposed into sub-bands, and energy-ratio criteria are used to select informative components for reconstruction. The reconstructed signals are subsequently fused into fixed-layout two-dimensional images via a sector-allocation strategy. Next, a diagnosis-aware quality assessment module assigns a quality score to each image to guide training. Low-quality samples are further regularized via cross-quality feature alignment using a quality-weighted supervised pairwise loss, encouraging them to align with high-quality counterparts from the same class. Finally, a Swin Transformer backbone performs classification. Experiments on the proprietary RGFD dataset and the public SEUB dataset demonstrate consistent gains under controlled mixed-quality settings, with accuracies of up to 99.4% and 98.7%, respectively, while additional evaluations under reproducible degradations further confirm the robustness of the proposed quality-aware learning mechanism. Zhuowei Li 0011, Jianxu Mao, Yaonan Wang 0001, Junfei Yi, Caiping Liu, Hui Zhang 0023 |
IEEE Trans. Reliab. | 4 |
| 2026 | PDE-Based Adaptive Consensus Control of Leader-Follower Multiagent Systems With Dynamic Event-Triggered StrategyabstractThis article addresses the leader–follower consensus problem for a class of nonlinear multiagent systems (MASs) whose collective behavior is modeled by a diffusion partial differential equation (PDE). Existing control strategies for such systems often suffer from high communication overhead and a lack of robustness to unknown nonlinearities and disturbances. To overcome these limitations, we introduce a novel adaptive control scheme that integrates a dynamic event-triggered mechanism with a radial basis function neural network (RBFNN) approximator. The dynamic event trigger scheme significantly reduces communication burdens by aperiodically updating the control signal only at specific moments, while the RBFNN is employed to effectively compensate for the unknown boundary function and unmodeled disturbances. We provide a rigorous Lyapunov-based stability analysis to prove that the proposed controller guarantees stability of the closed-loop system. Numerical simulations demonstrate the efficacy of the proposed method, showing a substantial reduction in communication frequency while ensuring precise consensus tracking. Zhongqi Lu, Yaonan Wang 0001, Zhiji Han, Zhijie Liu 0001, Hang Zhong, Wei He 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 2 |
| 2025 | Semantic Ambiguity Modeling and Propagation for Fine-Grained Visual Cross View Geo-LocalizationabstractVisual cross view geo-localization is generally approached within a joint retrieval-and-calibration framework. However, existing methods overlook semantic ambiguities arising from query and reference images characterized by low overlap, dynamic foregrounds, viewpoint changes, and perceptual aliasing. This makes it challenging to automatically control the relative importance of the two tasks, potentially compromising the retrieval task in favor of the offset regression. Consequently, the model may encounter conflicting dominating gradients during joint training. To address this, we propose to model the semantic ambiguity during the offset regression process by integrating associated uncertainty scores, represented as 2D Gaussian distributions, to mitigate negative transfer effects within the joint tasks. We further introduce an uncertainty-aware similarity metric to enhance similarity assessment between query and reference images, accounting for their semantic ambiguities. This metric propagates uncertainty scores into the retrieval task, focusing on certain samples and learning discriminative feature embeddings, allowing the model to adaptively handle conflicting dominating gradients during joint training. Extensive experiments demonstrate that our method improves the overall performance of the joint tasks, achieving state-of-the-art results on the VIGOR and CVACT datasets. Mingtao Feng, Fenghao Tian, Jianqiao Luo, Weisheng Dong, Yaonan Wang 0001, Ajmal Mian |
AAAI | 6 |
| 2025 | Feature Information Driven Position Gaussian Distribution Estimation for Tiny Object DetectionabstractTiny object detection remains challenging in spite of the success of generic detectors. The dramatic performance degradation of generic detectors on tiny objects is mainly due to the the weak representations of extremely limited pixels. To address this issue, we propose a plug-and-play architecture to enhance the extinguished regions. We for the first time exploit the regions to be enhanced from the perspective of pixel-wise amount of information. Specifically, we model the entire image pixels feature information by minimizing Information Entropy loss, generating an information map to attentively highlight weak activated regions in an unsupervised way. To effectively assist the above phase with more attention to tiny objects, we next introduce the Position Gaussian Distribution Map, explicitly modeled using a Gaussian Mixture distribution, where each Gaussian component's parameters depend on the position and size of object instance labels, serving as supervision for further feature enhancement. Taking the information map as prior knowledge guidance, we construct a multi-scale position gaussian distribution map prediction module, simultaneously modulating the information map and distribution map to focus on tiny objects during training. Extensive experiments on three public tiny object datasets demonstrate the superiority of our method over current state-of-the-art competitors. Jinghao Bian, Mingtao Feng, Weisheng Dong, Jianqiao Luo, Yaonan Wang 0001, Guangming Shi |
CVPR | 6 |
| 2025 | Cross-Modal Interactive Perception Network with Mamba for Lung Tumor Segmentation in PET-CT ImagesabstractLung cancer is a leading cause of cancer-related deaths globally. PET-CT is crucial for imaging lung tumors, providing essential metabolic and anatomical information, while it faces challenges such as poor image quality, motion artifacts, and complex tumor morphology. Deep learning-based models are expected to address these problems, however, existing small-scale and private datasets limit significant performance improvements for these methods. Hence, we introduce a large-scale PET-CT lung tumor segmentation dataset, termed PCLT20K, which comprises 21, 930 pairs of PET-CT images from 605 patients. Furthermore, we propose a cross-modal interactive perception network with Mamba (CIPA) for lung tumor segmentation in PET-CT images. Specifically, we design a channel-wise rectification module (CRM) that implements a channel state space block across multi-modal features to learn correlated representations and helps filter out modality-specific noise. A dynamic cross-modality interaction module (DCIM) is designed to effectively integrate position and context information, which employs PET images to learn regional position information and serves as a bridge to assist in modeling the relationships between local features of CT images. Extensive experiments on a comprehensive benchmark demonstrate the effectiveness of our CIPA compared to the current state-of-the-art segmentation methods. We hope our research can provide more exploration opportunities for medical image segmentation. The dataset and code are available at https://github.com/mj129/CIPA. Chenyu Lin, Yaonan Wang 0001, Hui Zhang 0023 |
CVPR | 4 |
| 2025 | Hierarchical Gaussian Mixture Model Splatting for Efficient and Part Controllable 3D Generationabstract3D content creation has achieved significant progress in terms of both quality and speed. Although current Gaussian Splatting-based methods can produce 3D objects within seconds, they are still limited by complex preprocessing or low controllability. In this paper, we introduce a novel framework designed to efficiently and controllably generate high-resolution 3D models from text prompts or images. Our key insights are three-fold: 1) Hierarchical Gaussian Mixture Model Splatting: We propose a hybrid hierarchical representation to extract fixed number of fine-grained Gaussians with multiscale details from textured object, also establish part-level representation of Gaussians primitives. 2) Mamba with adaptive tree topology: We present a diffusion mamba with tree-topology to adaptively generate Gaussians with disordered spatial structures, without the need for complex preprocessing and maintain linear complexity generation. 3) Controllable Generation: Building on the HGMM tree, we introduce a cascaded diffusion framework combining controllable implicit latent generation, which progressively generates condition-driven latents, and explicit splatting generation, which transforms latents into high-quality Gaussian primitives. Extensive experiments demonstrate the high fidelity and efficiency of our approach. Qitong Yang, Mingtao Feng, Weisheng Dong, Yaonan Wang 0001, Ajmal Mian |
CVPR | 6 |
| 2025 | Multi-range Adaptive Perception Transformer for Iterative Homography EstimationabstractHomography estimation is fundamental to various vision tasks. Iteration-based methods have recently achieved significant success in this field. However, errors introduced during iterations can lead to increased image deformation. Existing methods often focus on capturing local correspondences in the later stages of iteration while downplaying global ones, which may cause errors to persist and propagate into subsequent iterations, ultimately leading to error accumulation. To alleviate this issue, we propose Multi-range Adaptive Perception Transformer for Iterative Homography Estimation (MAPTHomo), which integrates Multi-range Attention (MRA) and Adaptive Perception Module (APM). Specifically, MRA captures both global and local correspondences, enabling the model to adapt to varying levels of deformation. The APM dynamically adjusts attention focus based on the current context. The combination of MRA and APM enhances the error-correction capability of the iterative process, effectively mitigating error accumulation. Extensive experiments demonstrate that MAPTHomo outperforms previous methods and exhibits strong generalization ability. Tianming Li, Qing Zhu 0003, Zhen Zhou 0003, Jianqiao Luo, Yaonan Wang 0001 |
ICASSP | 5 |
| 2025 | Prompt-Driven Transferable Adversarial Attack on Person Re-identification with Attribute-Aware Textual InversionabstractPerson re-identification (re-id) models are vital in security surveillance systems, requiring transferable adversarial attacks to explore the vulnerabilities of them. Recently, vision-language models (VLM) based attacks have shown superior transferability by attacking generalized image and textual features of VLM, but they lack comprehensive feature disruption due to the overemphasis on discriminative semantics in integral representation. In this paper, we introduce the Attribute-aware Prompt Attack (AP-Attack), a novel method that leverages VLM's image-text alignment capability to explicitly disrupt fine-grained semantic features of pedestrian images by destroying attribute-specific textual embeddings. To obtain personalized textual descriptions for individual attributes, textual inversion networks are designed to map pedestrian images to pseudo tokens that represent semantic embeddings, trained in the contrastive learning manner with images and a predefined prompt template that explicitly describes the pedestrian attributes. Inverted benign and adversarial fine-grained textual semantics facilitate attacker in effectively conducting thorough disruptions, enhancing the transferability of adversarial examples. Extensive experiments show that AP-Attack achieves state-of-the-art transferability, significantly outperforming previous methods by 22.9% on mean Drop Rate in cross-model&dataset attack scenarios. Yuan Bian 0002, Min Liu 0008, Yunqi Yi, Yaonan Wang 0001 |
ICCV | 6 |
| 2025 | Partially Matching Submap Helps: Uncertainty Modeling and Propagation for Text to Point Cloud Localization
Mingtao Feng, Longlong Mei, Jianqiao Luo, Fenghao Tian, Jie Feng 0003, Weisheng Dong, Yaonan Wang 0001 |
ICCV | 8 |
| 2025 | CVPT: Cross Visual Prompt Tuning
Lingyun Huang, Jianxu Mao, Junfei Yi, Ziming Tao, Yaonan Wang 0001 |
ICCV | 5 |
| 2025 | Wavelet-Based Distillation with Structured Frequency Alignment
Pengyu Lu, Junfei Yi, Jianxu Mao, Junlong Yu, Shuohao Xiao, Zhenyu He 0015, Yaonan Wang 0001 |
ICIG (2) | 7 |
| 2025 | Point Density Fusion for Multimodal 3D Object Detection
Ziyang Peng, Jianxu Mao, Wei He 0001, Caiping Liu, Zhenyu He 0015, Ziming Tao, Yaonan Wang 0001 |
ICIG (2) | 7 |
| 2025 | FFTA-Net: A Frequency-Domain Fusion and Temporal Alignment Network for Transmission Line Defect Detection
Jianxu Mao, Yaonan Wang 0001, Junlong Yu, Junfei Yi, Zhenyu He 0015, Ziming Tao, Hui Zhang 0023 |
ICIG (2) | 3 |
| 2025 | A Method for Constructing Building Structure Grid Map Based on a Climbing AlgorithmabstractAerial-terrestrial amphibious robots excel in search and rescue tasks in unstructured terrains but face challenges in autonomous navigation indoors. Traditional full-mapping methods can degrade global path planning performance, especially when semi-static obstacles shift, leading to suboptimal paths. We propose a method for constructing building structure grid maps that are unaffected by semistatic obstacles. Our approach includes a building structure recognition algorithm based on an octree structure to differentiate between occupied and free grid cells. Experimental results demonstrate that coverage path planning on building structure grid maps produces superior global paths compared to traditional grid maps, offering a more streamlined and robust solution for autonomous navigation of aerial-terrestrial amphibious robots in indoor environments. Xidong Zhou, Hang Zhong, Hui Zhang 0023, Yaonan Wang 0001 |
ICRA | 7 |
| 2025 | Tacit Learning with Adaptive Information Selection for Cooperative Multi-Agent Reinforcement Learning
Lunjun Liu, Weilai Jiang, Yaonan Wang 0001 |
AAMAS | 3 |
| 2025 | Foreground-Aware Enhancement-Based Multimodal 3D Object DetectionabstractLiDAR is one of the most widely used 3D detection sensors in applications such as autonomous driving and unmanned inspection. However, when uniformly sampling the entire scene to generate point cloud data, the number of foreground points reflected by target objects is often significantly lower than that of the background points, which have a larger coverage area. This imbalance poses considerable challenges to the performance of object detection models, especially in detecting small or distant objects. To overcome this challenge, this paper presents a Foreground-Aware Enhancement-based Multimodal 3D Object Detection Method (PFA), which effectively mitigates the low detection accuracy of small and distant objects caused by insufficient foreground points. The proposed method incorporates a Foreground-Aware Enhancement Module (FAEM) and a Region-Focused Attention Module (RFAM). The FAEM module enhances the model’s focus on foreground regions, while the RFAM module strengthens multimodal fused features. Together, these components significantly improve the detection accuracy of small and distant objects. Experimental results on the KITTI dataset demonstrate that the proposed method achieves 3D detection accuracies of 84.65%, 59.56%, and 71.48% for cars, pedestrians, and cyclists, respectively, under the hard evaluation level. Furthermore, the model also shows significant advantages on the KITTI public test set and validation set for both easy and moderate samples, fully validating its effectiveness and generalizability in enhancing multimodal 3D object detection accuracy. Ziyang Peng, Wei He 0001, Jianxu Mao, Ziming Tao, Junfei Yi, Yaonan Wang 0001 |
IJCNN | 6 |
| 2025 | Decentralized Multi-robot Navigation Policy with Enhanced Security Using Graph GRU Policy NetworkabstractFormulating a multi-robot obstacle avoidance policy is essential for enabling safe and efficient navigation in multi-robot environments, forming a critical component of the effective operation of multi-robot systems. Recently, reinforcement learning has been applied to improve the performance of decentralized, policy-driven robots in task execution. However, ensuring the safety of these agents during movement remains a significant challenge due to the inherent risks associated with the reinforcement learning process, such as frequent collisions. To address this issue and enhance the safety of policy-guided multi-robot navigation, we propose a novel policy based on imitation learning. This framework introduces a novel policy neural network that integrates a graph attention mechanism with the GRU network structure. The key innovation lies in utilizing the interactions between neighboring robots to enhance the safety of their movements. In a multi-robot simulation environment, robot behaviors are directed by the proposed policy. A comparative analysis was conducted between our approach and RL-RVO, one of the advanced methods in the field. The results demonstrate that our approach outperforms RL-RVO, achieving a higher success rate and significantly improving safety performance. Lin Chen 0034, Yuxuan Ao, Zhen Zhou 0003, Yaonan Wang 0001, Danwei Wang |
IROS | 4 |
| 2025 | Dual-Mode Passive Fault-Tolerant Control for Underwater Vehicles with Actuator Faults and Time-Varying DisturbancesabstractThis paper investigates the control problem of underwater vehicles subject to time-varying external disturbances and actuator faults. A novel passive fault-tolerant control (PFTC) scheme is developed to address the coupled disturbance-fault dynamics inherent in underwater vehicle systems. The proposed dual-mode architecture comprises: 1) a robust fault-tolerant control scheme based on high-order sliding mode observers (HOSMOs) for minor fault scenarios, which effectively compensates for bounded disturbances and partial actuator degradation; 2) a conditionally triggered estimation mechanism integrated with fault-tolerant control allocation (FTCA) and HOSMOs for severe fault conditions, enabling fault estimation and model compensation via event-triggered parameter updating. The hybrid architecture ensures computational efficiency by activating the estimation module only when predefined triggering conditions are violated. Comprehensive experimental results validate the superiority of the proposed method in maintaining stability and performance under various fault conditions. This work provides a systematic solution for underwater vehicle control under coupled disturbance-fault conditions, with verified real-time performance and implementation feasibility. Yizong Chen, Zhiqiang Miao, Kangcheng Liu, Yaonan Wang 0001 |
IROS | 5 |
| 2025 | Coordinated Energy-Trajectory Economic Model Predictive Control for Autonomous Surface Vehicles under DisturbancesabstractThe paper proposes a novel Economic Model Predictive Control (EMPC) scheme for Autonomous Surface Vehicles (ASVs) to simultaneously address path following accuracy and energy constraints under environmental disturbances. By formulating lateral deviations as energy-equivalent penalties in the cost function, our method enables explicit trade-offs between tracking precision and energy consumption. Furthermore, a motion-dependent decomposition technique is proposed to estimate terminal energy costs based on vehicle dynamics. Compared with the existing EMPC method, simulations with real-world ocean disturbance data demonstrate the controller’s energy consumption with a 0.06% energy increase while reducing cross-track errors by up to 18.61%. Field experiments conducted on an ASV equipped with an Intel N100 CPU in natural lake environments validate practical feasibility, achieving 0.22 m average cross-track error at nearly 1 m/s and 10 Hz control frequency. The proposed scheme provides a computationally tractable solution for ASVs operating under resource constraints. Zhongqi Deng, Yuan Wang 0040, Jian Huang 0001, Hui Zhang 0023, Yaonan Wang 0001 |
IROS | 5 |
| 2025 | Joint Optimization of Multi-Agent Task Allocation and Path Planning for Continuous Pickup and Delivery TasksabstractThe multi-agent pickup and delivery problem is central to coordinating multiple agents in real-world applications such as warehouse automation, urban logistics, and robotic delivery networks, where efficient task assignment and pathfinding are vital for maximizing production efficiency. However, existing approaches often struggle to seamlessly integrate task allocation with path planning while also failing to address the demands of continuous pickup and delivery tasks, resulting in suboptimal performance and limited scalability in dynamic environments. To address these problems, we first introduce a novel task allocation approach, which constructs a cost matrix to satisfy pickup and delivery timing constraints for tasks and employs a Mixed-Integer Linear Programming (MILP) model to compute a task assignment matrix queue. Next, the CBS-TAPF framework is proposed, which constructs search forests for tasks and paths to address the joint optimization of task allocation and path planning. This framework is further extended to Continuous Multi-Agent Pickup and Delivery (CMAPD) tasks by dynamically updating the task allocation matrix queue, enhancing robustness and adaptability for real-world, sustained scenarios. Finally, through simulation and real-world experiments, we validated the effectiveness of the proposed methods. The experimental results demonstrate its superiority across diverse environments, ensuring robust performance in various operational scenarios. Hongkai Fan, Bo Ouyang, Qinjing Xie, Yaonan Wang 0001, Zhi Yan 0002, Jiawen He, Qin Tan |
IROS | 4 |
| 2025 | A Parallel Fuzzy Nonlinear ADRC Framework for Robotic Machining with VibrationsabstractThis paper proposes a Parallel Fuzzy Nonlinear Active Disturbance Rejection Control (FNLADRC) strategy to improve the precision and robustness of robotic manipulators in machining large, complex components. By decoupling the multi-degree-of-freedom dynamics and integrating Nonlinear ADRC with fuzzy logic, the method adaptively estimates and compensates for disturbances in real time, effectively mitigating machining vibrations and trajectory errors. A parallel control architecture enhances multi-joint coordination, improving adaptability and response speed. Additionally, a fuzzy logic-based tuning mechanism dynamically adjusts control parameters to boost robustness. Simulations on a UR5 robotic arm validate the method’s superior performance in dynamic, uncertain environments, demonstrating FNLADRC as a promising solution for high-precision robotic machining. Xueqi Hao, Changqing Gao, Qiu Fang, Yaonan Wang 0001 |
IROS | 5 |
| 2025 | WLuav: An Air-Ground Robot with High Ground Adaptability and Trajectory Tracking PerformanceabstractAir-ground robots have received more and more attention and applications due to their air-to-ground motion performance and excellent energy efficiency. However, airground robots have many gaps including complex structure mechanisms, low terrain adaptability and low-precision controllers to significantly limit practical application. In this paper, an air-ground robot, wheel-leg unmanned aerial vehicle (WLuav), is proposed based on a five-link wheel leg structure to obtain excellent ground adaptive and air maneuvering capabilities. Based on the improved structure mechanism, a hierarchical adaptive agile controller is proposed to improve its trajectory tracking accuracy and ground adaptability. Besides, a mode switching strategy based on the support force solver is proposed to provide smooth and rapid mode switching. Finally, comprehensive experiments and a benchmark comparison are carried out to validate the performance of the proposed system, where the WLuav system shows excellent ground adaptive performance and trajectory tracking performance, and the energy efficiency can reach 79.46 %. Zhiqiang Miao, Chuanpeng Niu, Kangcheng Liu, Yaonan Wang 0001 |
IROS | 5 |
| 2025 | Optimization Based Human-Guided Variable-Stiffness Visual Impedance Control for Contact-Rich TasksabstractIn contact-rich tasks such as polishing and drilling, inevitable physical interactions often lead to task deviations due to interference, typically resulting in excessive contact forces and eventual task failure. To tackle these challenges, we propose an innovative human-guided visual-impedance control framework. Specifically, we first introduce an interaction model within image feature space, which models the dynamics of human-robot-environment interactions. Subsequently, human operation skills are characterized through human-guided wrenches, and acts on visual features through a projection matrix, thus integrating human-guided wrenches with visual-impedance interaction dynamics. Finally, leveraging this framework, we develop a novel variable-stiffness visual-impedance control strategy. The impedance parameters are optimized online via Quadratic Program, ensuring that the end-tool contact force converges to desired value while adhering to safety constraints. The validity of the proposed framework was established through polish experiments. Jiao Jiang, Yaonan Wang 0001, Yiming Jiang 0001, Danping Zeng, Chao Zeng 0002, Chenguang Yang 0001, Hui Zhang 0023 |
IROS | 2 |
| 2025 | MRMT-PR: A Multi-Scale Reverse-View Mamba-Transformer for LiDAR Place RecognitionabstractPlace recognition is a fundamental technology of high relevance for autonomous robot navigation. Existing methods encounter significant challenges arising from scene variations (e.g., illumination changes, dynamic objects), view-point shifts, and difficulties in data fusion and alignment. These factors often lead to a substantial drop in recognition recall, which is typically addressed in the literature by training deep neural networks to learn invariant feature representations. In this paper, we propose MRMT-PR, a novel multi-scale reverse-view Mamba-Transformer architecture for LiDAR-based place recognition that uses a single-frame point cloud as its input. Our MRMT-PR framework consists of a multi-scale reverse-view preprocessing module for LiDAR point clouds, a Mamba-Transformer feature encoder, and a global feature fusion module. This architecture effectively mitigates the impact of perspective and illumination variations, enhances the global representational capacity of LiDAR features, and significantly improves recognition robustness under challenging conditions such as viewpoint changes and long-term localization. Experiments conducted on NCLT dataset with challenging scenarios demonstrate that MRMT-PR outperforms existing LiDAR-based place recognition baselines in terms of overall performance. Jingwen Wang 0009, Hongshan Yu, Yaonan Wang 0001, Javier Civera 0001, Xieyuanli Chen |
IROS | 4 |
| 2025 | Efficient Multimodal 3D Object Detector via Instance-Level Contrastive DistillationabstractMultimodal 3D object detectors leverage the strengths of both geometry-aware LiDAR point clouds and semantically rich RGB images to enhance detection performance. However, the inherent heterogeneity between these modalities, including unbalanced convergence and modal misalignment, poses significant challenges. Meanwhile, the large size of the detection-oriented feature also constrains existing fusion strategies to capture long-range dependencies for the 3D detection tasks. In this work, we introduce a fast yet effective multimodal 3D object detector, incorporating our proposed Instance-level Contrastive Distillation (ICD) framework and Cross Linear Attention Fusion Module (CLFM). ICD aligns instance-level image features with LiDAR representations through object-aware contrastive distillation, ensuring fine-grained cross-modal consistency. Meanwhile, CLFM presents an efficient and scalable fusion strategy that enhances cross-modal global interactions within sizable multimodal BEV features. Extensive experiments on the KITTI and nuScenes 3D object detection benchmarks demonstrate the effectiveness of our methods. Notably, our 3D object detector outperforms state-of-the-art (SOTA) methods while achieving superior efficiency. The implementation of our method has been released as open-source at: https://github.com/nubot-nudt/ICD-Fusion. Zhuoqun Su, Huimin Lu 0002, Shuaifeng Jiao, Junhao Xiao 0001, Yaonan Wang 0001, Xieyuanli Chen |
IROS | 5 |
| 2025 | Safety-Aware Geometric Force-Impedance Control for ManipulatorsabstractSince its inception, impedance control has emerged as a fundamental framework for robotic interaction control. Recent advancements in geometric impedance control have demonstrated certain advantages over traditional Cartesian impedance control. However, existing geometric impedance control approaches generally lack force regulation capabilities or rigorous stability guarantees. In this paper, we propose a safety-aware geometric force-impedance controller that addresses these limitations. By incorporating an energy tank mechanism, the proposed approach enables precise force tracking while preserving full compatibility with the impedance behavior. Furthermore, an energy injection and freezing mechanism is introduced, allowing dynamic regulation of energy exchange between the tank and the robotic system. Notably, the proposed method eliminates the need for an offline estimation of the initial energy stored in the tank, facilitating real-time adjustments of force controller parameters. To validate the effectiveness of the proposed framework, we conduct extensive polishing experiments on a real robotic platform. The results demonstrate the capability of the proposed controller to achieve stable and precise force regulation. Danping Zeng, Yaonan Wang 0001, Yiming Jiang 0001, Jiao Jiang, Chenguang Yang 0001, Hui Zhang 0023 |
IROS | 2 |
| 2025 | GeoScene: Temporal 3D Semantic Scene Completion with Geometric Correlation between ImagesabstractSemantic Scene Completion (SSC) aims to reconstruct the entire 3D scene in terms of both occupancy and semantics, serving as a fundamental task for autonomous driving and robotic systems. Camera-based methods have seen significant advancements due to their low cost and rich visual cues. However, previous approaches have predominantly focused on semantic recovery. This can lead to inaccurate occupancy predictions and, consequently, the failure of downstream tasks such as trajectory planning. To address this limitation, we propose a novel multi-frame matching framework, GeoScene, which reconstructs spatial structures through inter-frame geometric correlations of temporal images and subsequently infers scene semantic information. Specifically, we extract features from distinct frames in the depth dimension and derive depth features by constructing a cost volume. Following this, dot product and voxelization operations are applied between the extracted features and depth features to correct assignment errors. Furthermore, we introduce a surface normal-based regression loss to preserve fine-grained surface structures. Extensive experiments on the SemanticKITTI dataset demonstrate that GeoScene outperforms existing state-of-the-art methods. Xiaogang Zhang 0002, Hua Chen 0008, Zhiqiang Miao, Yaonan Wang 0001, Kangcheng Liu |
IROS | 5 |
| 2025 | Dual Enhancement on 3D Vision-Language Perception for Monocular 3D Visual GroundingabstractMonocular 3D visual grounding is a novel task that aims to locate 3D objects in RGB images using text descriptions with explicit geometry information. Despite the inclusion of geometry details in the text, we observe that the text embeddings are sensitive to the magnitude of numerical values but largely ignore the associated measurement units. For example, simply equidistant mapping the length with unit 'meters' to 'decimeters' or 'centimeters' leads to severe performance degradation, even though the physical length remains equivalent. This observation signifies the weak 3D comprehension of pre-trained language model, which generates misguiding text features to hinder 3D perception. Therefore, we propose to enhance the 3D perception of model on text embeddings and geometry features with two simple and effective methods. Firstly, we introduce a pre-processing method named 3D-text Enhancement (3DTE), which enhances the comprehension of mapping relationships between different units by augmenting the diversity of distance descriptors in text queries. Next, we propose a Text-Guided Geometry Enhancement (TGE) module to further enhance the 3D-text information by projecting the basic text features into geometrically consistent space. These 3D-enhanced text features are then leveraged to precisely guide the attention of geometry features. We evaluate the proposed method through extensive comparisons and ablation studies on the Mono3DRefer dataset. Experimental results demonstrate substantial improvements over previous methods, achieving new state-of-the-art results with a notable accuracy gain of 11.94% in the 'Far' scenario. Our code will be made publicly available. Min Liu 0008, Yuan Bian 0002, Zhaoyang Li 0011, Gen Li 0008, Yaonan Wang 0001 |
ACM Multimedia | 7 |
| 2025 | NaME: A Natural Micro-expression Dataset for Micro-expression Recognition in the WildabstractMicro-expressions (MEs) are involuntary facial expressions that reveal genuine emotions and have significant applications in fields such as psychology, security, and human-computer interaction. However, previous ME datasets are mainly collected in controlled laboratory environments, such as fixed views, single illumination and head movements, limited subjects and the lack of background. There are significant gaps between them and the real world. To handle this issue, we introduce a novel Natural Micro-Expression (NaME) dataset, a natural dataset collected under unconstrained real-world conditions. It encompasses (1) diverse subjects, multiple views and varying head movements ; (2) rich background information, providing a more realistic benchmark for the micro-expression recognition (MER) research. Furthermore, we propose a MER benchmark for natural environments, named MixFormer. MixFormer includes an efficient sparse attention mechanism to capture subtle facial motions from various factors, and a face-background mix of attention module to model the environment context to help MER. Extensive experiments are conducted to analyze our NaME dataset and benchmark. We believe that our dataset and benchmark will pave the way for future research in MER beyond controlled settings, facilitating the deployment of MER in practical applications. NaME is available at github.com/real-ljt/NAMEdataset. Jiateng Liu, Hengcan Shi, Haiwen Liang, Yuan Zong, Yaonan Wang 0001, Wenming Zheng |
ACM Multimedia | 6 |
| 2025 | Decoupled Identity and Attribute Tokenization for Person Re-IdentificationabstractVision-language models like CLIP have revolutionized person re-identification (ReID) by enabling cross-modal semantic alignment. However, most of the existing CLIP-based ReID methods suffer from a critical limitation: semantic entanglement, where identity and attribute features are indiscriminately compressed into a single, undifferentiated token representation. This oversight fails to account for their inherently distinct roles in characterizing individuals.To address this limitation, we propose an Identity-Attribute-Decoupled Tokenization (IADT) method, a hierarchical framework with two synergistic components:Subject-oriented tokens that model identity through a cross-modality feature inverse mapping paradigm, preserving invariant biometric features;Attribute-aware tokens that capture localized characteristics through the cross-interaction of local features and learnable prototype vectors, dynamically focusing on discriminative regions without manual supervision.The hierarchical tokenization enables disentangled yet complementary representation learning: Identity and attribute semantics are encoded into distinct embedding subspaces, while cross-token contrastive learning establishes semantic reinforcement through attention-guided feature interaction. Crucially, this process does not require part-level annotations, making it directly applicable to real-world deployment. Extensive experiments validate effectiveness of the proposed method. For example, on the Market-1501 dataset, IADT achieves 97.1% mAP (+2.5% over SOTA) and 98.2% Rank-1 accuracy. For the challenging MSMT benchmark, it attains 88.9% mAP (+1.7% improvement) with 93.1% Rank-1 accuracy, demonstrating consistent superiority. The code will be available at https://github.com/llraay/IADT. Min Liu 0008, Yuan Bian 0002, Yaonan Wang 0001 |
ACM Multimedia | 5 |
| 2025 | Generalizing to New Area: Self-Distillation Curriculum Learning for Fine-Grained Cross View LocalizationabstractFine-grained cross-view localization seeks to predict ground-level camera positions within GPS-tagged aerial images by matching ground and aerial views. Existing methods often rely on large-scale ground truth annotations from specific regions, but performance degrades due to domain shifts when models trained in one area are applied to another. However, collecting region-specific annotations for each area is costly or infeasible. To address this, we propose a self-distillation curriculum learning framework that generalizes pretrained localization models to unseen new areas. Our approach introduces a Dirichlet-based quality assessment strategy to evaluate teacher-generated pseudo labels, where high uncertainty signals noisy predictions and low uncertainty indicates clean samples. This uncertainty is used to guide an easy-to-hard curriculum learning strategy, where easy samples are prioritized initially, and more challenging samples are progressively incorporated, enabling effective student training. Furthermore, we develop a joint optimization scheme that updates both the student model and pseudo labels, applying adaptive label smoothing to mitigate label noises and taking full advantage of new area data. Extensive experimental results on the VIGOR and KITTI benchmarks demonstrate that our method outperforms state-of-the-art approaches in new area localization, achieving superior accuracy without additional supervision. Fenghao Tian, Mingtao Feng, Jianqiao Luo, Longlong Mei, Weisheng Dong, Yaonan Wang 0001 |
ACM Multimedia | 8 |
| 2025 | Searching Efficient Semantic Segmentation Architectures via Dynamic Path SelectionabstractExisting NAS methods for semantic segmentation typically apply uniform optimization to all candidate networks (paths) within a one-shot supernet. However, the concurrent existence of both promising and suboptimal paths often results in inefficient weight updates and gradient conflicts. This issue is particularly severe in semantic segmentation due to its complex multi-branch architectures and large search space, which further degrade the supernet's ability to accurately evaluate individual paths and identify high-quality candidates. To address this issue, we propose Dynamic Path Selection (DPS), a selective training strategy that leverages multiple performance proxies to guide path optimization. DPS follows a stage-wise paradigm, where each phase emphasizes a different objective: early stages prioritize convergence, the middle stage focuses on expressiveness, and the final stage emphasizes a balanced combination of expressiveness and generalization. At each stage, paths are selected based on these criteria, concentrating optimization efforts on promising paths, thus facilitating targeted and efficient model updates. Additionally, DPS integrates a dynamic stage scheduler and a diversity-driven exploration strategy, which jointly enable adaptive stage transitions and maintain structural diversity among selected paths. Extensive experiments demonstrate that, under the same search space, DPS can discover efficient models with strong generalization and superior performance. Yuxi Liu 0019, Min Liu 0008, Yaonan Wang 0001 |
NeurIPS | 5 |
| 2025 | DRLHomo: Disentangled Representation Learning for Cross-Modal Homography Estimation
Tianming Li, Zhen Zhou 0003, Qing Zhu 0003, Jianqiao Luo, Yaonan Wang 0001 |
PRCV (9) | 5 |
| 2025 | EA-OSPGB: Multiple robots dynamic online algorithm for solving full coverage path planning of multiple robots in unknown terrain environments
Fangfang Zhang 0004, Jianbin Xin, Jinzhu Peng, Yaonan Wang 0001 |
Expert Syst. Appl. | 5 |
| 2025 | LRMM: Low rank multi-scale multi-modal fusion for person re-identification based on RGB-NI-TI
Di Wu 0046, Shenglong Gan, Qin Wan 0001, Yaonan Wang 0001 |
Expert Syst. Appl. | 7 |
| 2025 | Data-driven adaptive formation control based on preview mechanism for networked multi-robot systems with communication delays
Chenzhuolei Chao, Haoran Tan, Xueming Zhang, Gang Wang 0014, You Wu 0005, Yaonan Wang 0001 |
Neurocomputing | 6 |
| 2025 | Focus DETR: Focus detection transformer for ship wall-climbing robot real-time object detection
Xiaofang Yuan, Haozhi Xu, Yaonan Wang 0001 |
Neurocomputing | 5 |
| 2025 | EMPViT: Efficient multi-path vision transformer for security risks detection in power distribution network
Xiaofang Yuan, Haozhi Xu, Yaonan Wang 0001 |
Neurocomputing | 5 |
| 2025 | State-of-Health Prediction for Lithium-Ion Batteries Based on SIREN-TransformerabstractState-of-Health (SOH) is a critical indicator reflecting the degradation status of lithium-ion batteries (LIBs), which is essential for ensuring operational safety, prolonging lifespan, and optimizing battery management strategies. Consequently, accurate SOH prediction is paramount for effective battery health management and timely maintenance interventions. However, the primary difficulty no longer lies in achieving peak accuracy on a single benchmark but in maintaining consistent performance across heterogeneous scenarios. In real-world applications, battery usage patterns are often complex, highly dynamic, and nonstationary, which results in degradation trajectories that contain both subtle high-frequency fluctuations and long-term temporal trends. To address these challenges, this paper introduces a novel hybrid deep learning framework, termed the SIREN-Transformer model, which integrates the Sinusoidal Representation Network (SIREN) with the Transformer-based architecture. The SIREN module employs periodic activation functions to extract high-frequency features and capture subtle nonlinear patterns. Meanwhile, the Transformer component extracts long-range temporal relationships and contextual dependencies within the time-series data, enhancing the model’s capacity to generalize across diverse operational scenarios. The SIREN-Transformer model combines frequency-aware representation and long-range temporal modeling to accurately capture nonlinear, high-frequency, and complex degradation patterns in battery SOH prediction. Extensive experiments on multiple LIB datasets demonstrate the prediction accuracy and robustness, which validate the effectiveness of the proposed model for SOH prediction. Zeyan Sun, Xudong Wang 0008, Zhaoke Ning, Yaonan Wang 0001 |
IEEE Internet Things J. | 5 |
| 2025 | Hyperrectangle Embedding for Debiased 3D Scene Graph Prediction From RGB Sequencesabstract3D scene graph has emerged as a powerful high-level representation of the environment and is regarded as a prerequisite for long-term autonomous robotic operations. A practical research problem here is to predict the 3D scene graph from sequentially captured data. However, existing methods neglect the polysemy of semantic roles that coarse feature vectors are insufficient to represent entities in different relationship semantics. This extremely limits their capability to predict relationships. We propose an approach to tackle the aforementioned challenge by introducing a novel representation, the hyperrectangle embedding, which represents entity using distinctive geometry for more effective scene understanding, rather than learning within vector-based feature with blindly increasing dimensions. By incorporating an entity within two affine-transformed embeddings, each representing either the subject or object and characterized by separate learnable transformations, we achieve the polysemy of semantic roles. The intersections of affine-transformed hyperrectangle embeddings represent the bidirectional relationship between two entities. We identify bias and reliability as two challenges impeding the model learning process. In response to the bias, that arises from long-tailed distributions in the data, we propose a history-guided debiasing strategy that utilizes a confusion history block comprised of previous hyperrectangle embeddings. This strategy mitigates inherent biases by extracting pertinent information and facilitating knowledge transfer from dominant categories to rare ones. To enhance the reliability of predictions, we introduce predictive uncertainty into the 3D scene graph prediction task. We develop a post-hoc reliability enhancement strategy to identify potentially unreliable predictions and subsequently enhance the model's predictive accuracy. Extensive experiments on the 3DSSG dataset show the effectiveness of the proposed method in this challenging task, outperforming existing state-of-the-art. Mingtao Feng, Chenbo Yan, Weisheng Dong, Yaonan Wang 0001, Ajmal Mian |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2025 | Towards Real-World Aerial Vision Guidance With Categorical 6D Pose TrackerabstractTracking the object 6-DoF pose is crucial for various downstream robot tasks and real-world applications. In this paper, we investigate the real-world robot task of aerial vision guidance for aerial robotics manipulation, utilizing category-level 6-DoF pose tracking. Aerial conditions inevitably introduce special challenges, such as rapid viewpoint changes in pitch and roll and inter-frame differences. To support these challenges in task, we first introduce a robust category-level 6-DoF pose tracker (Robust6DoF). This tracker leverages shape and temporal prior knowledge to explore optimal inter-frame keypoint pairs, generated under a priori structural adaptive supervision in a coarse-to-fine manner. Notably, our Robust6DoF employs a Spatial-Temporal Augmentation module to deal with the problems of the inter-frame differences and intra-class shape variations through both temporal dynamic filtering and shape-similarity filtering. We further present a Pose-Aware Discrete Servo strategy (PAD-Servo), serving as a decoupling approach to implement the final aerial vision guidance task. It contains two servo action policies to better accommodate the structural properties of aerial robotics manipulation. Exhaustive experiments on four well-known public benchmarks demonstrate the superiority of our Robust6DoF. Real-world tests directly verify that our Robust6DoF along with PAD-Servo can be readily used in real-world aerial robotic applications. The project homepage is released at Robust6DoF. Yaonan Wang 0001, Danwei Wang |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2025 | NAS-PED: Neural Architecture Search for Pedestrian DetectionabstractPedestrian detection currently suffers from two issues in crowded scenes: occlusion and dense boundary prediction, making it still challenging in complex real-world scenarios. In recent years, Convolutional Neural Networks (CNN) and Vision Transformers (ViT) have shown their superiorities in addressing these issues, where ViTs capture global feature dependency to infer occlusion parts and CNNs make accurate dense predictions by local detailed features. Nevertheless, limited by the narrow receptive field, CNNs fail to infer occlusion parts, while ViTs tend to ignore local features that are vital to distinguish different pedestrians in the crowd. Therefore, it is essential to combine the advantages of CNN and ViT for pedestrian detection. However, manually designing a specific CNN and ViT hybrid network requires enormous time and resources for trial and error. To address this issue, we propose the first Neural Architecture Search (NAS) framework specifically designed for pedestrian detection named NAS-PED, which automatically designs an appropriate CNNs and ViTs hybrid backbone for the crowded pedestrian detection task. Specifically, we formulate transformers and convolutions with various kernel sizes in the same format, which provides an unconstrained space for diverse hybrid network search. Furthermore, to search for a suitable backbone, we propose an information bottleneck based NAS objective function, which treats the process of NAS as an information extraction process, preserving relevant information and suppressing redundant information from the dense pedestrians in crowd scenes Extensive experiments on CrowdHuman, CityPersons and EuroCity Persons datasets demonstrate the effectiveness of the proposed method. Our NAS-PED obtains absolute gains of 4.0% MR and 1.9% AP over the state-of-the-art (SOTA) pedestrian detection framework on CrowdHuman datasets. For the CityPersons and EuroCity Persons datasets, the searched backbone achieves stable improvement across all three subsets, outperforming some large language-image pre-trained models. Code will be released after acceptance. Min Liu 0008, Baopu Li, Yaonan Wang 0001, Wanli Ouyang |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2025 | DAUNet: A deformable aggregation UNet for multi-organ 3D medical image segmentation
Qinghao Liu, Min Liu 0008, Yuehao Zhu, Licheng Liu, Zhe Zhang 0022, Yaonan Wang 0001 |
Pattern Recognit. Lett. | 6 |
| 2025 | RF-REN: RGB-Frequency Relation Exploration Network for Micro-Expression RecognitionabstractMicro-expression recognition (MER) has drawn increasing attention in recent years due to its ability to reveal the true feelings people want to hide. The key challenge in MER is subtle motions, which are hard to capture but crucial for MER. Existing methods usually solve this problem by magnifying all motions in the whole face and temporal sequence. However, micro-expressions (MEs) only involve a few facial areas and several temporal snippets. The all-motion magnification in previous methods cannot precisely capture these local ME motion patterns, and can easily cause spatial as well as temporal distortions, which significantly decrease the MER accuracy. In this paper, we propose an RGB-Frequency Relation Exploration Network (RF-REN), which enhances the subtle motions in refined local ME cues by exploring spatial and temporal relations in both RGB and frequency domains. Specifically, we first decompose the ME video into RGB as well as frequency domains, and conduct temporal division according to different motion stages to cover various ME local patterns. Secondly, we construct an adaptive local-global relation exploration (LGRE) module to explore the local relation cues in the spatial appearance and temporal dynamics in both domains. Finally, we propose an RGB-Frequency routing strategy to fuse the RGB and frequency cues, aiming to aggregate spatial-temporal local-global information and enhance subtle motions for MER. Extensive experiments on three databases (CASME II, SAMM and SMIC) show that the proposed model outperforms other state-of-the-art methods. Jiateng Liu, Hengcan Shi, Yaonan Wang 0001, Yuan Zong |
IEEE Signal Process. Lett. | 3 |
| 2025 | HQCC: A Hybrid Quantum-Classical Classifier With Adaptive StructureabstractParameterized Quantum Circuits (PQCs) with fixed structures severely degrade the performance of Quantum Machine Learning (QML). To address this, a Hybrid Quantum-Classical Classifier (HQCC) is proposed. It adaptively optimizes the PQC through a dynamic circuit generator driven by Long Short Term Memory (LSTM) and exploits architectural plasticity to balance the entanglement performance and expressiveness, opening up a practical pathway for efficiently deploying QML in the Noisy Intermediate-Scale Quantum (NISQ) era. Extensive experiments on MNIST and Fashion-MNIST demonstrate that HQCC achieves up to 99.86% accuracy in binary classification and 97.12% in multi-class tasks, surpassing state-of-the-art quantum and classical baselines. Ren-Xin Zhao, Harun Siljak, Yunze He, Yaonan Wang 0001 |
IEEE Signal Process. Lett. | 5 |
| 2025 | STR: Spatial-Temporal RetNet for Distributed Multi-Robot NavigationabstractThe core of multi-robot collision avoidance is to guide robots to avoid collisions with other robots and obstacles in a dynamic multi-robot environment, which has recently gained increasing interest among the main challenges of robotics. However, the current multi-robot navigation policy neural network exhibits weak position encoding capabilities for spatial environmental features in mapping environment states and robot actions, as well as an inability to recurrently infer information on dynamic environmental features in the temporal dimension, leading to insufficient safety and effectiveness in guiding robot motion. In this paper, we propose a novel spatial-temporal RetNet (STR) that encodes reciprocal collision avoidance states between robots in both spatial and temporal dimensions, aiming to enhance the safety and effectiveness of the policy neural network in guiding robots to accomplish specified tasks. The spatial state encoder module is developed based on parallel RetNet structure, which enhances the ability of the neural network in multi-robot navigation policies to extract reciprocal collision avoidance states between robots in spatial dimensions and overcomes the weak position encoding capability of advanced transformer-based multi-robot navigation policy neural networks. A temporal state encoder is designed by introducing the recurrent RetNet structure. This enhances the multi-robot navigation policy neural network’s ability to encode features in the temporal dimension of multi-robot movements and overcomes the transformer-based multi-robot navigation policy neural network’s inability to recurrently infer information in the time dimension. Simulation experiments were designed to demonstrate that the safety and effectiveness of our proposed method outperform the previous state-of-the-art approaches in guiding the robot to complete the task. Physical experiments illustrate that our policy can be effectively applied to real-world systemsNote to Practitioners—Multi-robot navigation has a wide range of real-world applications, such as multi-robot formation flying for search and rescue, autonomous warehouse operations, and robots navigating through human crowds. This paper introduces a novel Spatial-Temporal RetNet (STR) framework aimed at enhancing safety and effectiveness in multi-robot collision avoidance. STR addresses the limitations of existing methods by improving the neural network’s ability to extract reciprocal collision avoidance states in both spatial and temporal dimensions. The spatial state encoder strengthens the extraction of spatial features, while the temporal state encoder improves the handling of time-dependent information. Simulation and physical experiments demonstrate that STR enhances robot navigation in dynamic environments, making it suitable for real-world applications such as multi-robot coordination. Lin Chen 0034, Yaonan Wang 0001, Zhiqiang Miao, Mingtao Feng, Yuanzhe Wang, Yang Mo, Wei He 0001, Hesheng Wang 0001, Danwei Wang |
IEEE Trans Autom. Sci. Eng. | 2 |
| 2025 | DIBNN: A Dual-Improved-BNN Based Algorithm for Multi-Robot Cooperative Area Search in Complex Obstacle EnvironmentsabstractAiming at the area search task of a multi-robot system in an unknown complex obstacle environment, we propose a cooperative area search algorithm based on a dual improved bio-inspired neural network (DIBNN). First, we improve the BNN model to reduce the interference of the complex obstacle environment on robot decision making. Each robot generally chooses the neuron with the largest sum of surrounding activity values among adjacent neurons as its next movement position. Then, we propose a collaborative search mechanism. When a robot falls into a local deadlock state in the complex obstacle environment, the mechanism will guide the robot to quickly find unsearched areas. Finally, we conduct multi-robot area search simulation experiments under different obstacle environments and compare them with three baseline algorithms in this field. The simulation results verify that the proposed algorithm can efficiently guide the multi-robot to complete the area search task in the complex obstacle environment.Note to Practitioners—The motivation of this article arises from the need to develop fast and effective area search algorithms for practical applications such as UAV swarm reconnaissance and multiple mobile robots area search and rescue. The algorithms based on BNN has been widely used in search tasks under unknown environments due to its good scalability and efficiency. However, the efficiency of area search in complex obstacle environments cannot be guaranteed. In order to achieve efficient area search in unknown complex obstacle environments, the DIBNN algorithm is proposed. It utilizes a cooperative search mechanism and achieves better performance. DIBNN can also be applied to multi-robot systems in different scenarios, demonstrating strong scalability. Bo Chen 0047, Hui Zhang 0023, Fangfang Zhang 0004, Yiming Jiang 0001, Zhiqiang Miao, Hongnian Yu, Yaonan Wang 0001 |
IEEE Trans Autom. Sci. Eng. | 7 |
| 2025 | A Lifelong Multi-Shuttle Scheduling Framework for the AS/RS SystemabstractEfficient scheduling of multi-shuttle is crucial for optimizing the performance of Automated Storage/Retrieval Systems (AS/RS). Multi-Agent Path Finding (MAPF) techniques play a pivotal role in addressing scheduling challenges by guaranteeing collision-free paths for multiple agents, however the assumptions of MAPF cannot meet the practical applications, such as 1) agents move at constant speeds and change directions instantaneously, 2) agents remain stationary at their destinations, 3) agents have different destinations. In this paper, we propose a novel lifelong multi-shuttle scheduling framework (LMSSF) to fill this gap based on the AS/RS system implemented in Hunan, China. LMSSF can perform re-planning and execution occur simultaneously, ensuring robust performance even in dynamic and uncertain environment. To apply MAPF to practical scenarios, we propose task conflict resolution, an improved single-shot MAPF and a path constraint mechanism to ensure collision-free movement of shuttles and improve the throughput performance of AS/RS system. Empirical evaluations on throughput and occupancy rate of outbound stations demonstrate the superiority of our proposed algorithm. Finally, LMSSF is applied to a real-world system and the experimental results show a 31.8% improvement in throughput compared to conventional strategies.Note to Practitioners—The motivation of this article stems from the challenges posed by the assumptions inherent in most MAPF algorithms when applied to practical AS/RS environments. Traditional MAPF algorithms typically compute discrete collision-free paths based on predefined start and goal locations. However, in real-world scenarios, tasks are continually generated, shuttles are highly dynamic, and various environmental constraints exist. To address these challenges, we propose the LMSSF as a means to adapt existing MAPF algorithms for practical AS/RS systems. This framework incorporates task conflict resolution, path planning optimization, and efficient shuttle control mechanisms. The paths obtained from the MAPF solver can be seamlessly integrated into our framework with minor format conversions. Through comprehensive testing, our algorithm demonstrates significantly improved throughput performance compared to traditional strategies. Hongkai Fan, Bo Ouyang, Zhi Yan 0002, Jiawen He, Zuozhi Zhang, Yaonan Wang 0001 |
IEEE Trans Autom. Sci. Eng. | 7 |
| 2025 | Trajectory-Attracted Adaptive Tracking Control for Robotic Systems Based on a Hybrid Guiding Vector Field in Flexible EnvironmentsabstractThis paper proposes a trajectory-attracted adaptive tracking control (TAATC) scheme based on a hybrid guiding vector field (HGVF) for robotic systems, addressing the high operational difficulty and safety concerns inherent in flexible environments. The HGVF is constructed using the characteristics of different task spaces in flexible environments to enable smooth transitions between free space and contact space with adjustable operating velocity, thereby enhancing robustness against environmental interactions and uncertainties. The HGVF can attract the state trajectories of the robotic systems by path convergence of an auxiliary dynamic system, simplifying the trajectory planning process with a time-independent representation of the desired path. In addition, a neural network is employed to compensate for uncertainties in the robotic systems, while a nonlinear duffing function represents the dynamic contact force model of flexible environments. In this way, the TAATC scheme is constructed by using the HGVF and the adaptive neural network, which unifies trajectory planning and tracking control. By using the proposed TAATC scheme, the robotic systems can achieve smooth interaction with flexible environments and obtain the desired tracking force without complex trajectory planning. The stability of the HGVF and the whole control scheme are analyzed by using the Lyapunov theorem. Finally, the effectiveness of the proposed method is validated through both simulation and experimental tests. Penghui Fan, Jinzhu Peng, Shuai Ding 0007, Yaoyu Yang, Yaonan Wang 0001 |
IEEE Trans Autom. Sci. Eng. | 7 |
| 2025 | Event-Triggered Super-Twisting Fixed-Time Consensus Control for Networked Nonlinear Multi-Agent Systems With DisturbanceabstractIn this paper, the leader-follower fixed-time consensus problem of networked nonlinear multi-agent systems (NNMASs) with unknown disturbance is investigated. A new event-triggered super-twisting fixed-time sliding mode consensus control (ESFSMCC) method is proposed, enabling all agents to achieve consensus in a fixed time. Firstly, a terminal sliding mode variable is designed to eliminate the convergence time dependence on the initial values of the system and avoid singularity. Secondly, a distributed event-triggered mechanism is developed to effectively reduce inter-agent communication frequency in a networked environment. In contrast to the existing fixed-time consensus control method, the improved super-twisting algorithm is adopted to construct a control protocol to achieve both global fixed-time consensus robustness and alleviate the issue of high-frequency chattering. Thirdly, the Lyapunov theory is utilized without the piecewise sliding mode technique to derive sufficient conditions for establishing the fixed-time stability of NNMASs, which still avoids the singularity problem. In this paper, the non-segmented terminal sliding mode is employed to prove the global fixed-time stability of the system states, thereby avoiding computational complexity. Finally, the effectiveness and advantages of the proposed method are verified through numerical simulations. Note to Practitioners—With the flourishing development of networks today, networked control is poised to become a future research hotspot, especially when large-scale agents interact and collaborate. The efficient utilization of network resources has become increasingly urgent due to the proliferation of such interactions. Consequently, reducing energy consumption poses a significant challenge. Moreover, the high-frequency chattering of control inputs presents a hindrance to the practical application of SMC. To address these issue, this paper proposes an event-triggered super-twisting distributed control protocol for NNMASs. This protocol not only eliminates the influence of initial values on stability time, thereby enhancing its practical value, but also mitigates the impact of input chattering, providing robust support for practical applications. Finally, a dynamics model of a multi-robotic manipulator is employed to verify the correctness and effectiveness of the proposed method. Additionally, underwater autonomous vehicle formations are expected to become valuable tools for future ocean resource exploration, while drone formations will likely become the preferred choice for agricultural development, geological exploration, and transportation. Even space exploration will involve coordinated interaction between multiple spacecraft. These are typical applications of NNMASs. Haoran Tan, Yun Feng 0001, Yaonan Wang 0001 |
IEEE Trans Autom. Sci. Eng. | 4 |
| 2025 | MQKIN: Manufacturing Quality Knowledge-Driven Interpretable Fault Diagnosis Network for Robotic Grinding EquipmentabstractDespite of the fast development of deep learning networks, its inexplicability poses a low credibility challenge for fault diagnosis methods based on them. This article proposes an interpretable fault diagnosis network for robotic grinding equipment driven by the knowledge of grinding process. The network consists of a vibration imaging module, a grinding quality quantification module, an auxiliary learning module, and a fault diagnosis module. In addition, the developed synchronous algorithm is embedded in the network to update the weights and coefficient matrix combinations, which can reconstruct the vibration images from the wavelet domain to be recognized by convolution kernels more easily. Besides, through auxiliary learning, the grinding knowledge can be learned by the network to endow the results with interpretability. Finally, the comparative experiment shows that the proposed method has a maximum accuracy of 4.25%, 7.25%, 2.25%, 3.25%, and 1.25% higher than the five popular state-of-the-art fault diagnosis algorithms, respectively; the ablation experiment shows that the proposed algorithm has a maximum accuracy of 6.25%, 12.5%, and 16.25% higher than each sub-algorithm, respectively; by analyzing the learned weights of the network, it can be concluded that the proposed network has successfully learned the potential features of grinding knowledge with satisfactory performance.Note to Practitioners—In the autonomous robotic grinding system, the reliability of robotic grinding equipment is a prerequisite for ensuring high-quality grinding of thin-walled parts for large equipment such as aircraft, high-speed rail, and ships. In robotic grinding manufacturing, the information of grinding quality is the most direct feedback of valuable information from the manufacturing system. Based on the causal relationship between the fault status of grinding equipment and the grinding quality of workpiece, this article develops an interpretable fault diagnosis network incorporated with the grinding knowledge, which treats grinding quality information as physical knowledge and provides credibility to the fault diagnosis results. Therefore, this article improves the interpretability of the fault diagnosis model in practical applications through the credibility of data and grinding knowledge, which improves the applicable potential of fault diagnosis methods based on deep learning in industry. Jianxu Mao, Yaonan Wang 0001, Zhe Li 0050, Xudong Wang 0008, Shaoyuan Wang |
IEEE Trans Autom. Sci. Eng. | 3 |
| 2025 | L₁ Adaptive Control-Based Formation Tracking of Multiple Quadrotors Without Linear Velocity Feedback Under Unknown DisturbancesabstractThis paper addresses the problem of formation control for a quadrotor swarm (QS) system with directed graph topology under external environmental disturbances and unreliable internal state acquisition. The proposed distributed robust control framework, based on a gemetric controller, incorporates${\mathcal {L}}_{1}$adaptive controllers and differentiator systems. First, the geometric formation controller is designed to implement the formation control of the nominal system. Then,${\mathcal {L}}_{1}$adaptive controllers are designed separately for each quadrotor’s position loop and attitude loop subsystems to address the effects of uncertainties such as external time-varying disturbances (matched and unmatched disturbances) and different mass variations of quadrotors. Furthermore, the differentiator system is devised to accurately estimate the higher-order derivatives of the non-directly-measurable velocity information and the virtual translation control signal, which enhances system accuracy while reducing computational complexity. The Lyapunov stability theory is employed to analyze the stability of the closed-loop system. Finally, the effectiveness and exceptional performance of this approach in QS formation control were validated through numerical simulation and experimental results. Note to Practitioners—The inspiration for this article comes from the issue of formation control in a cluster of quadrotor drones, which is also applicable to formation control in other types of drones. In this paper, a formation control algorithm based on${\mathcal {L}}_{1}$adaptive control strategy and arbitrary-order differentiation is designed. This algorithm can address not only the issue of time-varying wind disturbances frequently encountered during quadrotor drone flights but also the effects of unpredictable velocities and inconsistent masses of quadrotor drones. The disturbance rejection capability of this scheme enables quadrotor drones to be applied more safely and reliably in complex environments for search and rescue missions and surveillance tasks. Eliminating the need for linear velocity measurements reduces sensor costs and enhances system reliability and stability. The proposed formation control scheme allows the QS system to have different masses for each UAV, which can be applied to tasks such as collaboration logistics transportation, material delivery and crop spraying. Preliminary physical experiments have validated the feasibility of the proposed scheme, although it has not been applied in practical scenarios yet. In future research, we intend to equip each drone in the QS system with objects of different masses to achieve collaboration material transportation and delivery in complex environments. Zhiqiang Miao, Yaonan Wang 0001, Haoming Tang, Xiangke Wang, Wei He 0001 |
IEEE Trans Autom. Sci. Eng. | 3 |
| 2025 | A Novel Guided Deep Reinforcement Learning Tracking Control Strategy for MultirotorsabstractThis paper presents an intelligent control scheme for multirotors, where accurate trajectory tracking, strong robustness and reliable generalization are guaranteed by the dual-feedback sliding-mode (DFSM) guided deep reinforcement learning (RL). Different from current solutions, the proposed method explores optimal learning strategy on the sliding surface according to the DFSM demonstrations, where the elegantly designed parallel evaluation takes full advantage of model knowledge and learning exploration. Specifically, the intelligent tracking control is achieved in a two-step design. First, the DFSM algorithm is designed for multirotors, where the linear and nonlinear feedback terms work cooperatively. Second, the DFSM-guided deep RL is put forward to achieve intelligent switching on the sliding surface, where position and velocity errors are both considered to generate accurate switching decisions. In the framework, explorations and the DFSM demonstrations are evaluated in parallel, where only the explorations that are better than the DFSM baseline, are kept for policy improvement. In this way, the DFSM algorithm keeps pushing the RL policy to explore better strategy, where the unavoidable bad experiences arisen from exploration are identified accurately. Practical comparative experimental results are included to verify the effectiveness of the proposed strategy.Note to Practitioners—This paper is motivated by the practical problem of controlling multirotor system in uncertain environments. Up until now, most existing approaches are proposed without taking full advantage of model knowledge and deep learning techniques simultaneously, which lacks of reliability in practical application. To deal with the problem, a new dual-feedback sliding-mode (DFSM) guided deep reinforcement learning (RL) strategy is proposed, where the dual feedback and guided RL are designed to achieve satisfactory tracking control and simultaneously handle uncertainties. Specifically, by introducing double-check framework, the RL strategy explores optimal switching policy on the sliding surface according to the DFSM demonstrations, guaranteeing strong robustness and reliable generalization of the obtained RL policy in uncertain environments. The key feature of the framework is that the DFSM-driven training guarantees practice-oriented tracking control in a DFSM-RL cooperative manner. Comparative experiments are implemented to verify the tracking performance of the proposed intelligent control strategy. Hean Hua, Yaonan Wang 0001, Hang Zhong, Hui Zhang 0023, Yongchun Fang |
IEEE Trans Autom. Sci. Eng. | 2 |
| 2025 | Robust Adaptive Tracking Control for Aerial Transporting a Cable-Suspended Payload Using Backstepping Sliding Mode TechniquesabstractAerial transportation technology is the lifeline of air disaster rescue. In this article, a robust adaptive tracking control scheme using backstepping sliding mode techniques is proposed for a quadrotor-based aerial transportation system with a cable-suspended payload in disaster rescue, where the payload is ensured to be driven to predefined trajectories in the presence of strong coupling, uncertainties, and external disturbances. The quadrotor and the payload are modeled as a rigid body and a point mass, respectively, and the two coupling terms between the virtual input of the payload position loop and the payload attitude error as well as between the input force and the quadrotor attitude error are analyzed owing to the underactuated of the quadrotor-based transportation system. Then, adaptive backstepping sliding mode control strategies are designed for the position and swing dynamics of the payload to guarantee payload trajectory tracking, and an observer-based geometric attitude control method is presented for the quadrotor attitude dynamics to ensure the global attitude stability of the system, where prior information about disturbances is not required. The closed-loop stability of the whole system is strictly proven. Finally, real-world experiments are conducted to verify the feasibility and robustness of the proposed control scheme.Note to Practitioners—The motivation of this article is to investigate a robust and adaptive control tracking scheme for aerial transportation systems with a cable-suspended payload in disaster rescue. In most of the existing aerial transportation control schemes with a cable-suspended payload, the payload is driven to follow a desired trajectory while only considering the coupling effect between the aerial platform and the payload. However, in practical disaster rescue applications, the aerial transportation system is inevitably affected by strong coupling, uncertainties, and external disturbances. Meanwhile, to the authors’ best knowledge, there exist few studies that investigate payload following issues while considering strong coupling, uncertainties, and external disturbances simultaneously. Therefore, this article proposes a robust and adaptive tracking control scheme using backstepping sliding mode techniques for a quadrotor-based aerial transportation system with a cable-suspended payload to ensure the stable and accurate payload following control under strong coupling, uncertainties, and external disturbances, where prior information about disturbances is not required under the proposed scheme. The closed-loop stability of the whole system is strictly and mathematically analyzed as well as real-world experiments provide promising results. Moreover, the proposed scheme provides a more realistic setup for autonomous aerial transportation with cable-suspended supplies in disaster rescue. Jiacheng Liang, Yaonan Wang 0001, Hang Zhong, Hongwen Li, Hean Hua, Wei Wang 0025 |
IEEE Trans Autom. Sci. Eng. | 2 |
| 2025 | Self-Prior Guided Spatial and Fourier Transformer for Nighttime Flare RemovalabstractWhen capturing scenes with intense light sources, extensive flare artifacts often obscure the background and degrade image quality. Most flare removal methods directly process the flare-corrupted image as the optimization target, limiting the model’s understanding and generalization in complex real-world scenarios. In this paper, we propose a novel Self-prior Guided Spatial and Fourier Transformer (SGSFT) for nighttime flare removal. Specifically, we first establish a Self-prior Extraction Network to capture inherent priors in different scenes. Subsequently, we introduce a Semantic Contrast Enhancement Strategy to reinforce the semantic irrelevance between flare and light source, enabling the flare removal network to learn pattern differences between them and thus preserve light source. Finally, we build a Spatial and Frequency Flare Removal Network with Spatial Contextual Attention Block (SCAB) and Frequency Global Information Adjustment Block (FGIAB) to generate flare-free image. SCAB can perceive rich contextual information from self-prior guided regions and infer reasonable content. FGIAB captures global luminance representation in the frequency domain to maintain luminance consistency between the inferred regions and the flare-free areas. Extensive experiments demonstrate that the proposed approach achieves optimal performance in real nighttime scenes and exhibits robust generalization across various flare scenarios captured by different electronic devices. Note to Practitioners—The motivation of this paper is to remove flare artifacts in imaging. Flares degrade image quality and impact the performance of advanced vision tasks such as semantic segmentation and depth estimation in autonomous driving. Existing methods that indiscriminately extract contextual information from the entire image limit the model’s understanding of flares. This study proposes a self-prior guided flare removal network. The network first extracts self-prior information from flare-damaged images, then aggregates non-local information from the context indicated by the self-prior information to remove flares and infer semantically plausible fill content. Additionally, we model the global luminance information of the image in the frequency domain to enhance the global luminance consistency of the flare-free image. Experimental results show that our method has strong flare removal capabilities, but it also has a limitation. The training phase of this method requires paired flare-damaged images and flare images, which are difficult to obtain in real-world scenarios. Therefore, we will explore unsupervised flare removal methods in the future. Tianlei Ma, Zhiqiang Kai, Xikui Miao, Jing J. Liang, Jinzhu Peng, Yaonan Wang 0001, Hao Wang 0188, Xinhao Liu 0011 |
IEEE Trans Autom. Sci. Eng. | 6 |
| 2025 | Multi-Context Aggregation Network With Foreground Correction for Automated Few-Shot Defect SegmentationabstractState-of-the-art defect segmentation methods rely on sufficient training data and struggle to generalize to unseen categories. Few-Shot Semantic Segmentation (FSS) is introduced to specifically address these issues. However, existing FSS models still face two challenges in the industry. 1) Defects usually present as weak features, resulting in incomplete segmentation; 2) Severe background interference often leads to incorrect segmentation. To tackle these problems, we propose the Multi-Context Aggregation Network (MCANet). Specifically, we design a Cross-Layer Multi-Level Feature Aggregation Module (CMAM). CMAM effectively aggregates discretely distributed multi-level defect features across different layers and guides the query image to perceive defects from the pixel level, which avoids incomplete segmentation caused by weak features. Additionally, a Foreground Correction Module (FCM) is developed, which is equipped with a dedicated background predictor (BP) and a foreground corrector (FC). BP places more emphasis on learning features from backgrounds rather than defects. FC achieves efficient feature ensemble and further suppresses the backgrounds misidentified as defects in CMAM. They collaborate to prevent incorrect segmentation caused by background interference. Extensive experiments demonstrate the effectiveness of our method. We achieve state-of-the-art results on both FSSD-12, a public benchmark FSS dataset for strip steel, and FSS-AEB, an FSS dataset for aero-engine blades. Specifically, with 1/5 support images, we achieve 64.6%/65.6% mIoU on FSSD-12 and 55.0%/57.8% mIoU on FSS-AEB. Note to Practitioners—Surface defect segmentation has always been a hot topic in the industry. However, existing methods rely on sufficient training data and struggle to generalize to unseen categories, which significantly hinders the automation of defect segmentation. To address this problem, we propose MCANet for automated few-shot defect segmentation. It achieves effective segmentation for surface defects with limited data, even for unseen categories. Furthermore, MCANet achieves state-of-the-art results on two datasets from real-world industrial scenarios and also delivers significant improvements over the widely concerned large vision models. Finally, we integrate MCANet into an automated surface defect inspection platform consisting of an imaging system and a high-performance computing server for real-world performance validation. Yunfeng Ma, Min Liu 0008, Yuan Bian 0002, Yaonan Wang 0001 |
IEEE Trans Autom. Sci. Eng. | 6 |
| 2025 | SPDP-Net: A Semantic Prior Guided Defect Perception Network for Automated Aero-Engine Blades Surface Visual InspectionabstractAutomated surface defect detection is essential to manufacturing automation. However, automated inspection of aero-engine blades remains challenging due to tiny defects and weak features. To address this issue, we propose a semantic prior guided defect perception network, named SPDP-Net, which is ultimately integrated into an automated system to achieve efficient detection of defects. Firstly, a semantic prior mining module (SPM) is developed to capture finer-grained pixel-level location priors of defects by leveraging image feature mapping relations which facilitates the precise perception of tiny defects. Subsequently, we propose a defect enhancement perception module (DEP) to separate weak defects from complex backgrounds by utilizing defect location priors provided by SPM to enhance the features of defects while suppressing the values of non-defect regions, which makes the weak defects present as more obvious outliers. Finally, the global information extraction module (GIE) extracts the global features of defects, which helps to further improve the predicted results. When equipped with SPM, DEP and GIE, SPDP-Net can accurately identify and locate defects, exhibiting more competitive recognition and feature extraction capabilities for tiny defects and weak defects. To evaluate the effectiveness of our method, we construct an aero-engine blade surface defect detection dataset from real industrial scenarios called ABSDD with the collaboration of senior engineers. We achieve 95.9% precision, 94.0% recall and 94.9% F1 score on ABSDD. In addition, we also achieve state-of-the-art results on two public benchmark datasets, KSDD2 and DAGM. Finally, we have applied the developed SPDP-Net to an automated system and have conducted actual tests in collaboration with a well-known aero-engine production company.Note to Practitioners—At present, the automated detection system for surface defects in industrial manufacturing has not been well developed, which is particularly trailing behind in the field of aero-engine manufacturing. To the best of our knowledge, the surface defect detection of aero-engine blades is still carried out manually. To address this problem, we propose SPDP-Net for the automated detection of surface defects in aero-engine blades. It can accurately perceive and capture tiny and weak defects with excellent performance. SPDP-Net achieves state-of-the-art results on three tasks from different industrial manufacturing fields, which demonstrates its good transfer application capabilities. In addition, we integrate SPDP-Net into an automated detection system consisting of an autonomous imaging system and a high-performance computing server and conduct actual tests in an aero-engine production company. The test achieves remarkable results, indicating the good application prospects of the automated detection system. Yunfeng Ma, Min Liu 0008, Yiqiong Zhang, Yaonan Wang 0001 |
IEEE Trans Autom. Sci. Eng. | 5 |
| 2025 | Distributed Group Consensus Control for Multi-Agent Systems Based on Combined State InformationabstractThis paper addresses group consensus control for a class of multi-agent systems (MASs) by proposing a novel distributed group consensus protocol that leverages the combined state information (including both current and outdated states) of MASs. Unlike the standard protocol that uses either only current states or only outdated states, our approach integrates combined state information to achieve consensus. We derive sufficient conditions for the MASs to achieve group consensus, eliminating the dependence on two conservative assumptions present in previous literature. Furthermore, our proposed protocol improves the upper bound of input time delays compared to the standard approach. Several examples are provided to illustrate the effectiveness of our results. Note to Practitioners—The motivation behind this paper stems from the imperative to tackle the challenge of achieving group consensus control in multi-agent systems (MASs). Specifically, it focuses on the complexities and practical considerations inherent in MASs, such as the impact of competition and the influence of outdated states and input delays on system performance and stability. To address this issue, the paper introduces a novel group consensus protocol designed for second-order multi-agent systems that closely resemble real-world scenarios. This protocol integrates both position and velocity information as control inputs, enabling consensus under competitive conditions. By combining real-time and outdated states, the paper establishes conditions for achieving group consensus convergence while improving the bounds on input delays. Additionally, the paper extends the traditional bipartite graph representation to a directed graph, relaxing conservative assumptions and validating the effectiveness of the proposed protocol through theoretical analysis, numerical simulations, and comparative analysis with existing literature. Jiacheng Su, Yaonan Wang 0001 |
IEEE Trans Autom. Sci. Eng. | 4 |
| 2025 | Prototype, Modeling, and Control of Aerial Robots With Physical Interaction: A ReviewabstractThis article aims to investigate the research achievements related to aerial robots with physical interaction. Various morphologies of aerial physical interaction (APhI) robot prototypes with fixed wing, flapping wing, single main rotor, conventional underactuated multirotor, fully actuated multirotor, even deformed multirotor, and multiple platforms are reviewed for different APhI tasks associated with momentary, loose, and strong interaction coupling. This review also covers APhI robot rigid dynamics and robot-environment coupled interaction dynamics modeling methods, interaction wrench measurement/estimation, decoupled and coupled control, active aerial interaction control, and task-constrained planning approaches. Finally, future development directions and prospects are initially anticipated for aerial robots with physical interaction.Note to Practitioners—Aerial physical interaction (APhI) has been a hot topic in the field of aerial robots in recent years, which is a reflection of the advanced capabilities of aerial robots. However, APhI robots face challenges such as difficulty in flight stability and weak adaptability to dynamic environments while exerting active influence on environments. Under this background, this review aims to offer a reference for researchers and practitioners engaged in the related field from the aspects of system design, modeling, control, and task-constrained planning, which hopes to help them apply APhI robots to polar scientific expeditions, complex environment sampling, infrastructure inspection and maintenance, and other application areas. Further, this review also highlights the design idea of rigid-soft integrated APhI robots from the perspective of design-mechanism-performance to enhance interaction stability and safety. Hang Zhong, Jiacheng Liang, Hui Zhang 0023, Jianxu Mao, Yaonan Wang 0001 |
IEEE Trans Autom. Sci. Eng. | 6 |
| 2025 | Environment-Adaptive Motion Planning via Reinforcement Learning-Based Trajectory OptimizationabstractThis paper proposes a novel environment-adaptive motion planning framework for mobile robots, which utilizes deep reinforcement learning to dynamically adjust optimization objectives according to various environmental and robot-ego characteristics, greatly enhancing the adaptability and robustness compared to existing motion planning strategies. Our approach features a two-stage trajectory optimization algorithm that optimizes for smoothness, safety, and efficiency—elements critical in practical applications. Firstly, we propose a reinforcement learning algorithm that dynamically adjusts optimization objectives based on the environmental context, which carefully encodes the environment, coarse initial path and robot information into the observation space. Additionally, two techniques are designed to reduce the sim-to-real gap: 1) integrating classical optimization framework as the motion planning backbone; 2) using low-dimensional input in the learning component, which minimizes discrepancies between simulated and real-world conditions. With the hybrid strategy, not only the interpretability and stability of the classical motion planning pipeline is preserved, but also the adaptive capability of emerging DRL techniques is fully utilized. The efficacy of the proposed solution is demonstrated through extensive simulations and real-world experiments, showcasing superior performance in terms of safety and efficiency across various testing scenarios. Especially in unknown and cluttered real office environments, our approach significantly improve safety (17% increase in success rate) and efficiency (29% reduction in time cost), compared to the popular traditional methods. (Supplementary video link: https://youtu.be/ph0pDGpI864.). Zhejin Zhu, Runhua Wang, Yisong Wang 0001, Yaonan Wang 0001, Xuebo Zhang 0003 |
IEEE Trans Autom. Sci. Eng. | 4 |
| 2025 | History-Enhanced 3D Scene Graph Reasoning From RGB-D Sequencesabstract3D scene graph has emerged as a powerful high-level representation of the environment, and is considered a prerequisite for long-term autonomous robotic operations. However, building rich representations from RGB-D sequences remains a challenging problem. Existing methods ignore the semantic gap between linguistic and geometric feature spaces or neglect the importance of historical context in incrementally captured data. This limits the learning of visual-textual correspondence and the capability of relationship prediction. To address these problems, we propose a history-enhanced 3D scene graph reasoning framework that incrementally builds a consistent 3D semantic scene graph from an RGB-D image sequence. Specifically, we first introduce a cross-domain unified feature representation module to describe the object instances and their relationships distinctly. Next, we build a one-hot candidate matrix-enabled recurrent mechanism to reason the 3D scene graph, combining the perceived global and local history information. Finally, we design history-aware supervised semantics contrastive learning to optimize the scene-specific global history features. Extensive experiments on the 3DSSG dataset show the effectiveness of the proposed method in this challenging task, outperforming state-of-the-art approaches. Our code will be available athttps://github.com/cbyan1003/HE-3DSGR. Mingtao Feng, Chenbo Yan, Weisheng Dong, Yaonan Wang 0001, Ajmal Mian |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2025 | LDFCDet: Boosting 3D Object Detectors With Low-High Level Feature Crosses Using Laplace DistributionabstractHighly accurate 3D object detection is critical for autonomous driving and robotic sensing system. However, some objects with few foreground points significantly affect the accuracy of 3D object detection. As the network depth increases, the low-level features of these objects are gradually lost, especially for the hard object. Due to this issue, current LiDAR-only based and multimodal methods often misclassify background as foreground. Therefore, how to leverage the low-level feature that contain information about these objects in the high layer of the network becomes the key to optimizing the issue. In this paper, we propose LDFCDet, a framework boosting 3D object detectors with low-high level feature crosses using Laplace distribution(LD). In our proposed method, we design a low-high level feature crosses module(LHFCM) to embed low-level feature into high-level feature in the deeper layer of the network, and use Laplace distribution to obtain a new low-high level feature that includes information about these objects with few foreground points. In addition, we propose a res-gated feature aggregation module(RGFAM) to fuse the mutli-scale features. Our approach is well-suited for both LiDAR-based and multimodal methods.We evaluate the LDFCDet on the widely used KITTI dataset, and our method outperforms almost current 3D object detection methods on the challenging KITTI test set. Moreover, we conducted comparative experiments on the ONCE dataset, and the results further demonstrate the effectiveness and superiority of our method. Zhenyu He 0015, Jianxu Mao, Yaonan Wang 0001, Junlong Yu, Ziming Tao, Junfei Yi, Hui Zhang 0023, Shaoyuan Wang |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | RAMPGrasp: Retentive Attention-Based Multiscale Perception Grasp Detection NetworkabstractIn robotic grasp detection, challenges such as uncertainty in object type, size, and placement within the scene diminish grasping accuracy. However, the inability to effectively locate the graspable area and incomplete feature extraction for grasp detection are two key factors that hinder grasp detection accuracy and are not considered in current methods. This paper presents a novel retentive attention-based multiscale perception grasp detection network (RAMPGrasp) to address this constraint. First, we introduce retentive attention in the feature extraction module, which significantly improves the efficiency of attention score computation for long sequences in visual tasks. Second, we propose a multiscale spatial pyramid attention module, which can effectively adjust the importance of multiscale feature sequences and feature channels, while enhancing the correlation of multiscale features. Third, we design the prediction module as a coarse-to-fine framework, improving feature representation for grasp detection by considering the distribution trend of grasp poses. As a result, RAMPGrasp achieves state-of-the-art grasp detection accuracy, with 98.4% and 95.6% on the Cornell and Jacquard datasets, respectively. Jianan Huang 0002, Xuebing Liu, Qing Zhu 0003, Yaonan Wang 0001, Mingtao Feng, Zhen Zhou 0003, Lin Chen 0034, Danwei Wang |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | LFSSMam: Efficient Aggregation of Multi-Spatial-Angular-Modal Information Using Selective SSM for Light Field Semantic SegmentationabstractEfficiently aggregating 4D light field information to achieve accurate semantic segmentation has always faced challenges in capturing long range dependency information (CNN-based) and the memory limitations of quadratic computational complexity (Transformer-based). Recently, the Mamba architecture, which utilizes the state space model (SSM), has achieved high performance under linear complexity in various vision tasks. However, directly applying Mamba to 4D light field scanning will lead to an inherent loss of multi-spatial-angular information. To address the above challenges, we introduce LFSSMam, a novel Light Field Semantic Segmentation architecture based on the selective state space model (Mamba). Firstly, LFSSMam presents an innovative spatial-angular selective scanning mechanism to decouple and scan 4D multi-dimensional light field data. It separately captures the rich spatial context, complementary angular and structural information of light field 2D slices within the state space. In addition, we design an SSM-attention Cross-Fusion Enhance Module to perform preferential scanning and fusion across multi-spatial-angular-modal light field information, adaptively aggregating and enhancing the central view features. Comprehensive experiments on synthetic and real world datasets demonstrate that LFSSMam achieves leading edge SOTA (State-Of-The-Art) performance (with a 6.97% improvement to LF-based methods) while reducing memory and computational complexity. This work provides valuable guidance for the efficient modeling and application of multi-spatial-angular information in light field semantic segmentation. Our code is available at https://github.com/HNU-WQW/LFSSMam. Wenbin Yan, Hua Chen 0008, Qingwei Wu, Xiaogang Zhang 0002, Qiu Fang, Shengjie Hu, Yaonan Wang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2025 | Discriminative Correspondence Estimation for Unsupervised RGB-D Point Cloud RegistrationabstractPoint cloud registration is a fundamental task for estimating the rigid transformation matrix between two point clouds, and is regarded as a prerequisite for downstream vision tasks. Recent works have sought to address the registration problem using the obtainable RGB-D sequence, rather than relying solely on point clouds, which may not always be available. However, most existing unsupervised RGB-D point cloud registration works struggle to obtain fine-grained, robust, discriminative correspondences due to the simple concatenation of multimodal features and the increase in vector dimensions. These methods typically follow a common paradigm: extracting features from the input data, estimating correspondences, and obtaining the transformation matrix through geometric fitting. In this work, we design a generative feature extraction module to fully leverage multimodal information, and seek a novel perspective for correspondence estimation which expands the points in the source and target point clouds into hyperrectangle-based embeddings and considers their inner relationships, based on intersections in n-dimensional space, as the basis for estimating correspondences. Each hyperrectangle-based embedding is built upon the natural and discriminative semantics from the proposed generative feature extraction module, which involves a diffusion branch, a geometric branch, and point-pixel fusion. We harness the capability of the generative model to fully leverage the information from both complementary modalities in RGB-D frames. Furthermore, this distinctive geometry space allows for efficient calculation of intersection volumes and model conditional probabilistics for estimating correspondences. Extensive experiments on the 3DMatch and ScanNet datasets show the effectiveness of the proposed method in this challenging task, outperforming state-of-the-art approaches. Our code will be released at:https://github.com/cbyan1003/DCE. Chenbo Yan, Mingtao Feng, Yulan Guo, Weisheng Dong, Yaonan Wang 0001, Ajmal Mian |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2025 | FMSD: Focal Multi-Scale Shape-Feature Distillation Network for Small Fasteners Detection in Electric Power SceneabstractIn the electric power scene, fasteners play a pivotal role in securing and connecting electrical equipment, with small fastener detection (SFD) being crucial for ensuring operational stability. Despite the replacement of manual inspection methods by non-destructive techniques employing deep learning, these approaches often demand substantial computational resources and involve numerous parameters. While knowledge distillation (KD) can be a viable solution, existing KD methods may often fail to achieve satisfactory performance when dealing with small object presentation and little inter-class variability in SFD tasks. To alleviate this, we propose a Focal Multi-scale Shape-feature Distillation Network (FMSD) to achieve efficient and precise fastener detection in electric power scenarios. Specifically, we propose a novel Multi-Scale Shape-Aware Feature Aggregation module (MSFA) to augment the network's perception of object shape and scale during the KD process. Additionally, we propose a Contour-Guided Distillation (CGD) module to optimize the transfer of the extracted shape-sensitive knowledge between the teacher and student models. Through a series of experiments compared with existing state-of-the-art (SOTA) methods, our method demonstrates superior performance over existing SOTA techniques, both efficiently and effectively. Furthermore, validation on publicly available power scene datasets confirms the generalizability and adaptability of our proposed FMSD across various settings. Junfei Yi, Jianxu Mao, Hui Zhang 0023, Mingjie Li 0006, Kai Zeng 0010, Mingtao Feng, Xiaojun Chang, Yaonan Wang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 8 |
| 2025 | A Subspace Search-Based Evolutionary Algorithm for Large-Scale Constrained Multiobjective Optimization and ApplicationabstractLarge-scale constrained multiobjective optimization problems (LSCMOPs) exist widely in science and technology. LSCMOPs pose great challenges to algorithms due to the need to optimize multiple conflicting objectives and satisfy multiple constraints in a large search space. To better address such problems, this article proposes a dynamic subspace search-based evolutionary algorithm for solving LSCMOPs. The main idea is to initially allow the population to search in a low-dimensional subspace to increase convergence, then the searched subspace is gradually expanded to encourage the population to further search the full decision space. Specifically, the contribution of each decision variable to the evolution is first calculated using the proposed decision variable analysis method. Then, a probability-based offspring generation strategy is developed to encourage the population to preferentially search in a low-dimensional subspace composed of decision variables with high contribution degrees, thus speeding up the early convergence. With the continuous progress of evolution, the subspace is gradually expanded to ensure that the population can better explore the entire space. The performance of the proposed algorithm is evaluated on a variety of test problems with 100-1000 decision variables. Experimental results on four test suits and three real-world instances show that the proposed algorithm is efficient in solving LSCMOPs. Xuanxuan Ban, Jing J. Liang, Kunjie Yu, Kangjia Qiao, Ponnuthurai N. Suganthan, Yaonan Wang 0001 |
IEEE Trans. Cybern. | 6 |
| 2025 | A Novel Neural Dynamics Controller for Weakening the Chaos of Permanent Magnet Synchronous Generator and Its Extended ApplicationabstractThe permanent magnet synchronous generator (PMSG) system becomes unstable when unpredicted chaos appears, and current approaches do not take how to lessen this chaos phenomenon into account. Motivated by the ability of projective synchronization (PS) to adjust the chaotic system trajectory, this research aims to use PS to reduce the chaos in PMSG system. For better control in the time estimation of PS and the robustness of systems, an adaptive predefined-time robust zeroing neural dynamic controller (APTRZNDC) for the PS between PMSG systems is proposed. In the process, an adaptive parameter determined by the system error is designed with the demand for higher convergence factor in the case of large error. In addition, a nonlinear activation function contributed to the predefined-time synchronization is created, making the upper bound of synchronization time independent of system initial states and parameters, except for a single predefined parameter. Moreover, essential theorems for the predefined-time PS and robustness under the APTRZNDC are supplied and validated. And better robustness of PMSG system with the APTRZNDC is demonstrated when compared with other controllers. Furthermore, the APTRZNDC is applied in secure communication via the PS of PMSG systems, which guarantees both the timeliness of signals and the immunity of communication. Linju Li, Lin Xiao 0002, Qiuyue Zuo, Ping Tan 0004, Yaonan Wang 0001 |
IEEE Trans. Cybern. | 5 |
| 2025 | Adaptive Optimal Consensus Control for Nonlinear Uncertain Multiagent Systems Under DoS AttacksabstractThis article addresses the optimal control problem for nonlinear multiagent systems (MASs) with an uncertain nonlinear leader subject to intermittent Denial-of-Service (DoS) attacks. The main challenge is estimating the leader's dynamics when the uncertain nonlinear dynamics of the leader are unknown to all followers and communication between subsystems is intermittently disrupted by attacks. Furthermore, the uncertainty in the followers' dynamics adds complexity, making it difficult to eliminate reliance on the identifier network. To overcome these challenges, we develop a learning-based adaptive distributed observer to estimate the leader's dynamics under attacks. Based on this observer, a single-critic optimal consensus tracking control scheme is proposed to solve the leader-follower consensus problem in uncertain MASs without requiring an identifier network. It is proven that all system signals are uniformly ultimately bounded (UUB), and consensus tracking is achieved. The effectiveness of the proposed method is validated through a simulation example. Meijian Tan, Zhi Liu 0001, Ci Chen 0002, Yaonan Wang 0001, C. L. Philip Chen |
IEEE Trans. Cybern. | 4 |
| 2025 | Adaptive Safety-Based Tracking Control for Uncertain Robotic Systems With Input-Output Constraints: A Neural Network-Based Augmented High-Order Control Barrier Function ApproachabstractThis article investigates the trajectory tracking control of uncertain robotic systems with limited control torque input bounds and joint position constraints. A novel neural network-based augmented high-order control barrier function (NN-AHoCBF) is proposed to facilitate the tracking control strategy of uncertain robotic systems with input-output constraints, where the neural network (NN) is used to estimate uncertainties in the robotic system dynamics, and the bounds of NN approximation errors and NN weights are adapted in the high-order time derivative of the HoCBFs. The NN-AHoCBF is then derivated with a series of time-varying functions, and auxiliary systems are constructed to guarantee the time-varying functions to be HoCBFs. In this way, the control input of the robotic system is relaxed by adjusting the time-varying functions through the inputs of auxiliary systems in NN-AHoCBF barrier conditions. Also, the sufficient condition for the NN-AHoCBF is provided to adaptively ensure system safety. The adaptive safety-based tracking control method is designed based on NN-AHoCBF in quadratic program (QP) framework, which can not only satisfy input-output constraints simultaneously, but also achieve good robustness and tracking performance. A simulation example is performed on a two-DOF robotic mainpulator to verify the effectiveness of the developed controller. Haijing Wang, Jinzhu Peng, Wei He 0001, Yaonan Wang 0001 |
IEEE Trans. Cybern. | 5 |
| 2025 | An Interpretable Quantum Adjoint Convolutional Layer for Image ClassificationabstractThe interpretability of quantum machine learning (QML) refers to the capability to provide clear and understandable explanations for the predictions and decision-making processes of QML models. However, most quantum convolutional layers (QCLs) utilize closed-box structures that are inherently devoid of interpretability, leading to the opacity of principles and the suboptimal mapping of classical data. This significantly undermines the reliability of QML models. In addition, most of the current QML interpretability focuses on post hoc interpretability seriously neglecting the importance of exploring intrinsic causes. To tackle these challenges, we introduce the quantum adjoint convolution operation (QACO). It is an intrinsic interpretability scheme based on quantum evolution, as its quantum mapping precisely corresponds to the position and pixel values of the image and its principle is equivalent to the Frobenius inner product (FIP). Furthermore, we extend the QACO concept into the quantum adjoint convolutional layer (QACL) by integrating the quantum phase estimation (QPE) algorithm, enabling the parallel computation of all FIPs. Experimental results on PennyLane and TensorFlow platforms demonstrate that our method achieves a 6.3%, 3.4%, and 2.9% higher average test accuracy on Fashion MNIST, MNIST, and DermaMNIST datasets compared to classical and uninterpretable quantum counterparts, respectively, while maintaining 73.3% noise-robust accuracy under Gaussian noise, showcasing its superior generalizability and resilience in practical scenarios. Mengyi Wang 0003, Ren-Xin Zhao, Licheng Liu, Yaonan Wang 0001 |
IEEE Trans. Cybern. | 5 |
| 2025 | Task-Oriented Tool Manipulation With Robotic Dexterous Hands: A Knowledge Graph Approach From Fingers to FunctionalityabstractA primary challenge in robotic tool use is achieving precise manipulation with dexterous robotic hands to mimic human actions. It requires understanding human tool use and allocating specific functions to each robotic finger for fine control. Existing work has primarily focused on the overall grasping capabilities of robotic hands, often neglecting the functional allocation among individual fingers during object interaction. In response to this, we introduce a semantic knowledge-driven approach to distribute functions among fingers for tool manipulation. Central to this approach is the finger-to-function (F2F) knowledge graph, which captures human expertise in tool use and establishes relationships between tool attributes, tasks, and manipulation elements, including functional fingers, components, required force, and gestures. We also develop a manipulation element-oriented prediction algorithm using knowledge graph semantic embedding, enhancing the prediction of manipulation elements' speed and accuracy. Additionally, we propose the functionality-integrated adaptive force feedback manipulation (FAFM) module, which integrates manipulation elements with adaptive force feedback to achieve precise finger-level control. Our framework does not rely on extensive annotated data for supervision but utilizes semantic constraints from F2F to guide tool manipulation. The proposed method demonstrates superior performance and generalizability in real-world scenarios, achieving an 8% higher success rate in grasping and manipulation of representative tool instances compared to the existing state-of-the-art methods. The dataset and code are available at https://github.com/yangfan293/F2F. Fan Yang 0063, Wenrui Chen, Sijie Wu, Xin Li 0082, Zhiyong Li 0001, Yaonan Wang 0001 |
IEEE Trans. Cybern. | 7 |
| 2025 | A State Space Model for Multiobject Full 3-D Information Estimation From RGB-D ImagesabstractVisual understanding of 3-D objects is essential for robotic manipulation, autonomous navigation, and augmented reality. However, existing methods struggle to perform this task efficiently and accurately in an end-to-end manner. We propose a single-shot method based on the state space model (SSM) to predict the full 3-D information (pose, size, shape) of multiple 3-D objects from a single RGB-D image in an end-to-end manner. Our method first encodes long-range semantic information from RGB and depth images separately and then combines them into an integrated latent representation that is processed by a modified SSM to infer the full 3-D information in two separate task heads within a unified model. A heatmap/detection head predicts object centers, and a 3-D information head predicts a matrix detailing the pose, size and latent code of shape for each detected object. We also propose a shape autoencoder based on the SSM, which learns canonical shape codes derived from a large database of 3-D point cloud shapes. The end-to-end framework, modified SSM block and SSM-based shape autoencoder form major contributions of this work. Our design includes different scan strategies tailored to different input data representations, such as RGB-D images and point clouds. Extensive evaluations on the REAL275, CAMERA25, and Wild6D datasets show that our method achieves state-of-the-art performance. On the large-scale Wild6D dataset, our model significantly outperforms the nearest competitor, achieving 2.6% and 5.1% improvements on the IOU-50 and 5°10 cm metrics, respectively. Qing Zhu 0003, Yaonan Wang 0001, Mingtao Feng, Jian Liu 0014, Jianan Huang 0002, Ajmal Mian |
IEEE Trans. Cybern. | 3 |
| 2025 | Multiscale Spherical Feature Decoupling Network for Multimodal Image RegistrationabstractMultimodal image registration plays a crucial role in advancing Earth science. However, significant appearance variations and geometric deformations between multimodal images pose considerable challenges to this task. In this paper, we propose a novel multiscale spherical feature decoupling network (MSFDNet) for multimodal image registration by combining a multiscale iterative strategy with a multimodal decoupling strategy. MSFDNet adopts a multiscale architecture, with each scale incorporates a spherical feature decoupling (SFD) module with a carefully crafted three-stage decoupling strategy to bridge the modality gap. Specifically, we first introduce asymmetric shared and unique feature encoders to extract modality-shared and modality-unique features. Next, we design a spherical constraint learning (SCL) module to project the extracted features into spherical space, leveraging its regularized distance properties to enhance feature separability during the decoupling process. Finally, we propose a dual-path reconstruction mechanism that combines self-reconstruction with homography-guided cross-reconstruction to reconstruct the original multimodal images from the decoupled features, thereby simultaneously enhancing the learning of both feature decoupling and registration network. Based on the decoupled modality-shared features, we predict the registration function in a multiscale iterative manner, effectively bridging the geometric gap. Extensive experiments on multiple multimodal datasets validate the effectiveness of MSFDNet and demonstrate its state-of-the-art performance. Tianming Li, Zhen Zhou 0003, Qing Zhu 0003, Jianqiao Luo, Yaonan Wang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2025 | Uncertainty Guided Deep Lucas-Kanade Homography for Multimodal Image AlignmentabstractHomography estimation for multimodal images poses a considerable challenge in computer vision because of content disparities and the diverse feature points captured by different sensors. Existing methods typically extract feature maps using neural networks and apply the Lucas-Kanade (LK) algorithm, which is based on the brightness constancy assumption, to solve the homography matrix. However, applying this assumption across all pixel features in multimodal images can lead to inaccuracies, as these images often contain noise, such as homogeneous regions or considerable appearance variations, which can corrupt the network’s training. To address this problem, we propose an uncertainty-guided deep LK (UG-DLK) framework that integrates uncertainty predictions to enhance the network’s iterative learning process. Specifically, we employ a probabilistic approach where the network predicts the distribution of the feature map rather than fixed values. By designing an uncertainty neighborhood estimator, we unfold the cost volume along the channels into 2-D slices, allowing the model to focus on neighborhood information at specific locations, effectively reducing the interference from spatial neighborhoods in the estimation of feature uncertainty. Through uncertainty modeling, the network can accurately identify scenes and objects that comply with the brightness constancy constraint, leading to more robust learning outcomes. Additionally, we introduce a novel loss function that incorporates feature uncertainty, leading to a smoother optimization landscape near the true homography parameters and reducing convergence oscillations. Our method, which is evaluated on benchmark datasets such as Google Maps, Google Earth, MSCOCO, and DPDN, demonstrates state-of-the-art performance, confirming the robustness and adaptability of our model across various scenarios. Zhen Zhou 0003, Jianqiao Luo, Qing Zhu 0003, Yaonan Wang 0001, Hang Zhong, Mingtao Feng, Lin Chen 0034 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2025 | Modality Unified Attack for Omni-Modality Person Re-IdentificationabstractDeep learning based person re-identification (re-id) models have been widely employed in surveillance systems. Recent studies have demonstrated that black-box single-modality and cross-modality re-id models are vulnerable to adversarial examples (AEs), leaving the robustness of multi-modality re-id models unexplored. Due to the lack of knowledge about the specific type of model deployed in the target black-box surveillance system, we aim to generate modality unified AEs for omni-modality (single-, cross- and multi-modality) re-id models. Specifically, we propose a novel Modality Unified Attack method to train modality-specific adversarial generators to generate AEs that effectively attack different omni-modality models. A multi-modality model is adopted as the surrogate model, wherein the features of each modality are perturbed by metric disruption loss before fusion. To collapse the common features of omnimodality models, Cross Modality Simulated Disruption approach is introduced to mimic the cross-modality feature embeddings by intentionally feeding images to non-corresponding modality-specific subnetworks of the surrogate model. Moreover, Multi Modality Collaborative Disruption strategy is devised to facilitate the attacker to comprehensively corrupt the informative content of person images by leveraging a multi modality feature collaborative metric disruption loss. Extensive experiments show that our MUA method can effectively attack the omni-modality re-id models, achieving 55.9%, 24.4%, 49.0% and 62.7% mean mAP Drop Rate, respectively. Yuan Bian 0002, Min Liu 0008, Yunqi Yi, Yunfeng Ma, Yaonan Wang 0001 |
IEEE Trans. Inf. Forensics Secur. | 6 |
| 2025 | Learnable Prompts With Neighbor-Aware Correction for Text-Based Person Search
Min Liu 0008, Yaonan Wang 0001 |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2025 | Locally Aware Visual State Space for Small Defect Segmentation in Complex Component ImagesabstractSegmenting small defects within large imaging fields remains challenging in industrial scenarios due to the difficulty in distinguishing defects from complex component backgrounds and identifying defects comprising only a few pixels in high-resolution images. To address these issues, we propose a novel dual-branch feature extraction architecture, the locally aware visual state space block, which captures global contextual information while maintaining locally aware perception. In addition, we introduce the parallel quad-directional scanning fusion module to extract multiscale information, aggregating high-level features at different scales for enhanced global information fusion. To avoid losing small target details when upsampling the global segmentation mask to high-resolution input size, we develop progressive location refinement modules to incrementally refine small defect localization from the bottom up. Extensive experiments on our proposed small defect segmentation dataset and a public PCB dataset demonstrate that our method outperforms existing state-of-the-art methods in both performance and efficiency. Jinghao Bian, Mingtao Feng, Weisheng Dong, Jianqiao Luo, Yaonan Wang 0001, Guangming Shi |
IEEE Trans. Ind. Informatics | 7 |
| 2025 | VSLNet: Multimodal Data Fusion Network for Tree Species Classification in Overhead Transmission Line CorridorsabstractThe classification of tree species for overhead transmission lines (OHTL) is of great significance, facilitating the the timely removal of safety hazards posed by trees on power lines. Addressing the challenges in classifying OHTL line tree species, including subtle differences in target shape appearance, densely distributed targets, and limited representation in single-modal data, this article proposes a tree species classification network, VSLNet, based on multimodal data fusion. VSLNet constructs three asymmetric branches, which automatically select more discriminative features among spectra during spectral information processing, and jointly guide the extracted visible light information, ensuring global and local consistency for accurate multispectral classification. Furthermore, in LiDAR processing, the segmentation of individual trees contributes data such as tree height and crown diameter, and seamlessly integrates GPS data with multispectral classification results. Experimental results demonstrate that VSLNet is a feasible and reliable solution for tree classification, with potential applicability to other multimodal tasks. Hui Zhang 0023, Hang Zhong, Yihong Cao, Yaonan Wang 0001 |
IEEE Trans. Ind. Informatics | 6 |
| 2025 | CLMFNet: Cross-Level Multimodal Fusion Network for RGB-T Semantic Segmentation of Distribution Network LinesabstractAccurate semantic segmentation is crucial in distribution network line monitoring to ensure the system’s reliability and security. Due to the complexity of the environment and the diversity of devices, unimodal images (such as RGB images) struggle to provide enough information for effective segmentation. To address these challenges, the complementary nature of RGB and thermal infrared (TIR) images is leveraged to significantly enhance segmentation performance. Therefore, an innovative cross-level multimodal fusion network (CLMFNet) is proposed to improve the accuracy and robustness of semantic segmentation by integrating RGB and TIR data. A dual-branch architecture is employed in CLMFNet to extract features from both RGB and TIR images, which are then effectively integrated through a multimodal fusion strategy. Additionally, a cross-layer guidance mechanism is introduced to facilitate the complementation and optimization of features across different levels. CLMFNet was validated on a custom RGB-T dataset, and experimental results showed that it outperformed state-of-the-art methods in key metrics such as mean accuracy (mAcc) and mean intersection over union (mIoU), demonstrating its effectiveness in performing semantic segmentation in complex power distribution scenarios. Hui Zhang 0023, Hang Zhong, Yaonan Wang 0001 |
IEEE Trans. Ind. Informatics | 5 |
| 2025 | Deep Reinforcement Learning-Based Hierarchical Motion Planning Strategy for MultirotorsabstractThis article proposes a novel hierarchical motion planning strategy for multirotors, where the virtual goal (VG) oriented deep reinforcement learning (RL) and motion optimization are designed cooperatively to achieve efficient, flexible and smooth navigation in unknown environments. Specifically, the intelligent hierarchical motion planning is achieved in a three-step design. First, the dynamic VG generation algorithm is proposed considering the perception range of onboard sensors and current velocity, which transforms the global navigation into a real-time point-to-VG planning, thereby guaranteeing efficient computation even in resource-limited multirotors. Second, instead of generating motion actions, the upper-layer deep RL is designed to make spatial-temporal decisions of VG online, which outputs time allocation and spatial distribution commands according to current observation. Third, based on upper-layer's decisions, local optimization and control are implemented accordingly. Different from existing solutions, high-performance planning is guaranteed by the online VG oriented intelligent decision making, where the data-driven learning and model-driven optimization are integrated to navigate the multirotors. Comparative experiments are carried out in both physical simulation and indoor environments, which demonstrate the satisfactory performance of the proposed motion planning strategy in terms of feasibility, efficiency, navigation smoothness, and flexibility. Hean Hua, Yaonan Wang 0001, Hang Zhong, Hui Zhang 0023, Yongchun Fang |
IEEE Trans. Ind. Informatics | 2 |
| 2025 | Toward Efficient Power Scene Detection via Topology-Preserved Knowledge DistillationabstractThe power industry relies on efficient inspection systems to ensure stability and safety. While deep learning has advanced automated inspection, its reliance on custom modules for specific tasks can impact efficiency. Knowledge distillation (KD) offers a balanced solution, but the complex textures and structures of power equipment challenge conventional KD methods, which often fail to capture essential local semantic and topological relationships. To address this, we proposeTopNet, a novel topology-preserved KD framework for power scene detection tasks. Specifically, we model the teacher’s knowledge as a graph, where nodes encode local fine-grained features and edges capture global topological relationships. Based on this, we introduce node feature distillation and edge feature distillation to transfer local–global structural knowledge, which can enhance the student’s ability to perceive objects. Furthermore, we also introduce aggregated feature distillation to incorporate and transfer contextual semantic knowledge. Comprehensive experiments are conducted on two different benchmark datasets to demonstrate that TopNet achieves state-of-the-art detection performance with high efficiency, offering a robust solution for automated power equipment inspection. Junfei Yi, Tengfei Liu 0005, Jianxu Mao, Yaonan Wang 0001, Hui Zhang 0023, He Xie, Hang Zhong, Xiaojun Chang |
IEEE Trans. Ind. Informatics | 4 |
| 2025 | Balancing Accuracy and Efficiency With a Multiscale Uncertainty-Aware Knowledge-Based Network for Transmission Line InspectionabstractReal-world transmission line inspections (RTLIs) ensure power stability and safety. Deep learning (DL) models have become prevalent approaches for performing RTLI tasks. However, the high computational demands and substantial parameter requirements of DL models limit their real-world applicability. This article introduces a novel approach, a multiscale uncertainty-aware knowledge-based network, which is designed to balance the accuracy and efficiency in RTLI tasks. Specifically, we propose an uncertainty-aware knowledge distillation method that incorporates pixel-level uncertainty into the knowledge transfer process, mitigating the impact of noisy knowledge derived from extra background information contained in ground truths. In addition, our method integrates a multiscale relationship distillation technique, thus enhancing the transfer of multiscale information between the teacher and student models. Consequently, RTLI tasks can be efficiently accomplished using the well-learned lightweight student model. Comprehensive experiments conducted on a real-world dataset collected via uncrewedaerial vehicles demonstrate the efficacy of our proposed approach in terms of achieving high detection accuracy with reduced computational costs. Junfei Yi, Jianxu Mao, Hui Zhang 0023, Yurong Chen 0003, Tengfei Liu 0005, Kai Zeng 0010, He Xie, Yaonan Wang 0001 |
IEEE Trans. Ind. Informatics | 8 |
| 2025 | LBFormer: Scene Perception Segmentation Transformer Based on Local BlockabstractScene perception for autonomous vehicles and vessels is crucial for autonomous navigation. Current mainstream transformer methods typically split the feature map into windows, such as local, dilated, and horizontal/vertical bar windows. However, their token interaction is confined to fixed windows, posing challenges for image-based semantic segmentation. This article proposes a novel model, LBFormer, which enables flexible token interaction across different windows. Specifically, a window-level affinity graph is constructed from coarse-grained features using self-attention clustering and evolves during training, retaining top-k windows with high semantic relevance for each window. Self-attention purification is then employed to compress and filter fine-grained features with low semantic relevance within the top-k windows, ensuring effective token interaction for each feature point. To enhance context modeling within windows and build a more effective window-level affinity graph, a dual branch method extracts multidimensional features from each window, which are then interacted with and fused via the feature aggregation module. Extensive experiments at an image resolution of 224×224 were conducted on our private YZ-DATA water surface scene dataset and the public CamVid urban scene dataset. The results show that LBFormer achieves an MIoU of 89.80% on YZ-DATA and 61.31% on CamVid, surpassing mainstream transformer methods. Yunze He, Baoyuan Deng, Hongjin Wang, Liang Cheng 0005, Yaonan Wang 0001 |
IEEE Trans. Ind. Informatics | 8 |
| 2025 | Multimodal Fusion Network for Power Tower Semantic Segmentation and Inclination DetectionabstractAs the most fundamental supporting infrastructure of the distribution network, power towers require regular checks of their tilting status to ensure the system’s smooth operation. To overcome distinguishing objects in large-scale scenes using solely images or point clouds poses significant difficulties, we propose a multimodal fusion semantic segmentation network (MFSS) for power tower semantic segmentation and inclination detection. First, effective near-ground filtering and fixed-area slicing algorithms are proposed to address the issues of sample imbalance and insufficient data. Second, MFSS integrates RGB information and point cloud features to enhance the descriptive ability of the tower. Finally, a novel inclination detection method for distribution towers is proposed, estimating tower tilt from the axis between top and bottom centroids to improve accuracy and stability. Experimental results on our constructed dataset show that the proposed method outperforms existing algorithms, achieving 77.3% IoU and 96.69% per-class accuracy in tower segmentation. The mean angle deviation for tilt detection is 0.78$^{\circ }$, with a state judgment false rate of just 2.4%. Hui Zhang 0023, Hang Zhong, Youyuan Tang, Yihong Cao, Yaonan Wang 0001 |
IEEE Trans. Ind. Informatics | 8 |
| 2025 | An Adaptive Nearest Point Routing Method Based on Charging Nest Deployment Optimization for UAVs Power Tower Inspection
Hui Zhang 0023, Zhiwen Xu, Bo Chen 0047, Hean Hua, Hang Zhong, Wenhao Mo, Yaonan Wang 0001 |
IEEE Trans. Ind. Informatics | 7 |
| 2025 | Learning to Learn Transferable Generative Attack for Person Re-IdentificationabstractDeep learning-based person re-identification (re-id) models are widely employed in surveillance systems and inevitably inherit the vulnerability of deep networks to adversarial attacks. Existing attacks merely consider cross-dataset and cross-model transferability, ignoring the cross-test capability to perturb models trained in different domains. To powerfully examine the robustness of real-world re-id models, the Meta Transferable Generative Attack (MTGA) method is proposed, which adopts meta-learning optimization to promote the generative attacker producing highly transferable adversarial examples by learning comprehensively simulated transfer-based cross-model&dataset&test black-box meta attack tasks. Specifically, cross-model&dataset black-box attack tasks are first mimicked by selecting different re-id models and datasets for meta-train and meta-test attack processes. As different models may focus on different feature regions, the Perturbation Random Erasing module is further devised to prevent the attacker from learning to only corrupt model-specific features. To boost the attacker learning to possess cross-test transferability, the Normalization Mix strategy is introduced to imitate diverse feature embedding spaces by mixing multi-domain statistics of target models. Extensive experiments show the superiority of MTGA, especially in cross-model&dataset and cross-model&dataset&test attacks, our MTGA outperforms the SOTA methods by 20.0% and 11.3% on mean mAP drop rate, respectively. The source codes are available at https://github.com/yuanbianGit/MTGA. Yuan Bian 0002, Min Liu 0008, Yunfeng Ma, Yaonan Wang 0001 |
IEEE Trans. Image Process. | 5 |
| 2025 | LCTC: Lightweight Convolutional Thresholding Sparse Coding Network Prior for Compressive Hyperspectral ImagingabstractCompressive spectral imaging has garnered significant attention for its ability to effectively enhance the captured spatial and spectral information. Predominant methods, based on compressive sensing, typically formulate the imaging task as a constrained optimization problem and rely on hand-crafted priors to model the sparsity of spectral images. However, these approaches often suffer from suboptimal performance due to the inherent difficulty of identifying an appropriate transform space where spectral images exhibit sparsity. To overcome this limitation, we propose a novel convolutional sparse coding-inspired untrained network prior for fast and adaptive identification of the sparse transform domain and compressible signal. Specifically, a Lightweight Convolutional Thresholding sparse Coding (LCTC) network is designed as the sparse transform domain, with its inputs interpreted as sparse coefficients. Crucially, both the transform domain and its coefficients are solved in a self-supervised learning manner. Furthermore, we demonstrate that LCTC prior can be seamlessly incorporated into the iterative optimization algorithm as a Plug-and-Play (PnP) regularization. Both the LCTC and PnP-LCTC exhibit superior performance compared to previous methods. Experiments under various scenarios validate the effectiveness and efficiency of our approach. Yurong Chen 0003, Yaonan Wang 0001, Xiaodong Wang 0026, Xin Yuan 0002, Hui Zhang 0023 |
IEEE Trans. Image Process. | 2 |
| 2025 | Unsupervised Range-Nullspace Learning Prior for Multispectral Images ReconstructionabstractSnapshot Spectral Imaging (SSI) techniques, with the ability to capture both spectral and spatial information in a single exposure, have been found useful in a wide range of applications. SSI systems generally operate within the 'encoding-decoding' framework, leveraging the synergism of optical hardware and reconstruction algorithms. Typically, reconstructing desired spectral images from SSI measurements is an ill-posed and challenging problem. Existing studies utilize either model-based or deep learning-based methods, but both have their drawbacks. Model-based algorithms suffer from high computational costs, while supervised learning-based methods rely on large paired training data. In this paper, we propose a novel Unsupervised range-Nullspace learning (UnNull) prior for spectral image reconstruction. UnNull explicitly models the data via subspace decomposition, offering enhanced interpretability and generalization ability. Specifically, UnNull considers that the spectral images can be decomposed into the range and null subspaces. The features projected onto the range subspace are mainly low-frequency information, while features in the nullspace represent high-frequency information. Comprehensive multispectral demosaicing and reconstruction experiments demonstrate the superior performance of our proposed algorithm. Yurong Chen 0003, Yaonan Wang 0001, Hui Zhang 0023 |
IEEE Trans. Image Process. | 2 |
| 2025 | NPC-SPU: Nonlinear Phase Coding-Based Stereo Phase Unwrapping for Efficient 3D Measurementabstract3D imaging based on phase-shifting structured light is widely used in industrial measurement due to its non-contact nature. However, it typically requires a large number of additional images (multi-frequency heterodyne (M-FH) method) or introduces intensity features that compromise accuracy (space domain modulation phase-shifting (SDM-PS) method) for phase unwrapping, and it remains sensitive to motion. To overcome these issues, this article proposes a nonlinear phase coding-based stereo phase unwrapping (NPC-SPU) method that requires no additional patterns while maintaining measurement accuracy. In the encoding stage, a novel nonlinear distortion feature is introduced, while the signal-to-noise ratio of the phase codeword is preserved. In the decoding stage, a local phase unwrapping method that does not require additional auxiliary information is first proposed, closely associating the distortion information in the local wrapped phase. Then, a pre-calibrated stereo constraint system is used to filter potential matching phases, significantly reducing phase ambiguity and computational costs. Finally, to avoid the time-consuming and complex intensity kernel matching used in traditional methods, we propose a local phase correlation matching (LPCM) technique that enables lightweight and robust phase unwrapping. Experimental results demonstrate that this algorithm significantly enhances 3D reconstruction performance in scenarios with large depth, large disparity, complex colored structures, and dynamic scenes. Specifically, in dynamic environments (20mm/s), the proposed method achieves a lower measurement error rate (0.7829% vs. 6.4962%) with only 3 patterns, compared to the traditional three-frequency heterodyne (T-FH) method (using 9 patterns). Additionally, its measurement accuracy outperforms the advanced SDM-PS method, which also uses 3 patterns (0.1102 mm vs. 0.3232 mm). Ruiming Yu, Hongshan Yu, Wei Sun 0028, Yaonan Wang 0001, Naveed Akhtar, Kemao Qian |
IEEE Trans. Image Process. | 4 |
| 2025 | Robust Driving Intention Prediction Based on Multi-Stage Learning Under Vehicle-Infrastructure Cooperative PerceptionabstractIn mixed traffic of human-driven vehicles (HDVs) and connected and automated vehicles (CAVs), it is essential to predict the driving intention of HDVs to avoid potential risks. Data quality is crucial to intention prediction under a cooperative vehicle-infrastructure system, whereas the data collection of HDVs relies on the vehicle-infrastructure cooperative perception, which is inevitably exposed to perception errors. In this paper, a robust driving intention prediction framework based on multi-stage learning is proposed in mixed traffic under the vehicle-infrastructure cooperative perception situation. To address this issue, different information sources from vehicles and traffic are considered to derive the implied vehicle dynamic interaction relation and traffic flow context. A feature extraction module is developed to respectively capture the local and global features based on convolutional neural network (CNN) for reducing the impact of noise, which ensures the prominent detailed and overall descriptions of driving intention can be comprehensively acquired. Then, the deep multi-scale technique and multi-layer perceptron network are introduced to further extract deep features, and improve the model adaptability by complementary feature learning mechanism and nonlinear mapping ability. Experiment results on a real-world dataset confirm the effectiveness of our proposal in reducing the impact of poor data quality and accurately predicting driving intention. Xiaofang Yuan, Zhe Li 0050, Xiangcheng Pan, Yaonan Wang 0001 |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2025 | Adaptive Robust Formation Control Strategy for CAVs Merging at Highway On-RampsabstractHighway on-ramps are typical bottlenecks for the deployment of connected and automated vehicles (CAVs). However, few studies have considered time-varying traffic volume and dynamic uncertainties in real-world scenarios. This paper proposes an adaptive robust formation control strategy (ARFC) for CAVs to address these challenges. The ARFC consists of two stages: 1) In the tactical stage, the constraint-oriented platoon formation algorithm classifies approaching CAVs into multiple local virtual platoons (LVP), assigns the collision-free merging sequences and driving modes, thereby enabling the strategy to adapt to time-varying traffic volume. Unified spatial-dependent constraints are formulated to construct this geometry topology, where the spatial dependence ensures collision avoidance in the merging conflict zone. 2) In the operational stage, an adaptive robust control scheme is designed based on the Lyapunov min-max approach, rendering uniform boundedness, uniform ultimate boundedness performance of tracking errors, and string stability of LVP, regardless of time-varying uncertainties. Comprehensive validations demonstrate that the proposed strategy performs effectively under different traffic demands, and improves ride comfort and fuel economy compared to baseline methods. Xiangcheng Pan, Xiaofang Yuan, Zhigang Ling, Zhe Li 0050, Yaonan Wang 0001 |
IEEE Trans. Intell. Transp. Syst. | 6 |
| 2025 | Exploring Hierarchical Spatial Layout Cues for 3D Point Cloud Based Scene Graph Predictionabstract3D scene graph prediction is important for intelligent agents to gather information and perceive semantics of their environments. However, constructing an effective graph is nontrivial given the complexity of natural scenes. Existing solutions for graph representation of 3D scenes still distinguish each detailed discrepancy among all the relationships as flat thinking, ignoring the mechanism used by humans to perform this task. Inspired by the role of the prefrontal cortex in hierarchical reasoning, we analyze this problem from a novel perspective: exploring hierarchical spatial layout cues in 3D space and navigating that hierarchy to make the 3D scene graph more accurate in a vertical division to horizontal propagation strategy. To this end, we first encode the contextual object features for fine-gained object category classification. Next, we build a bottom-up hierarchical graph to predict remarkably diverse support relationships in a single concept regardless of numerous irrelevant relationships. Finally, equipped with the spatially-true and semantically-meaningful support relationships, we focus on the local region layout to propagate the semantic features to predict the additional non-support relationships under the guidance of the given referred hierarchical graph nodes. Experiments on the challenging 3DSSG benchmark show that our algorithm outperforms existing state-of-the-art, and can also alleviate the impact of the long-tailed distribution of training data. Our code is available athttps://github.com/HHrEtvP/HSLC-3DSG/. Mingtao Feng, Haoran Hou, Liang Zhang 0010, Yulan Guo, Hongshan Yu, Yaonan Wang 0001, Ajmal Mian |
IEEE Trans. Multim. | 6 |
| 2025 | Federated Hallucination Translation and Source-Free Regularization Adaptation in Decentralized Domain Adaptation for Foggy Scene UnderstandingabstractSemantic foggy scene understanding (SFSU) emerges a challenging task under out-of-domain distribution (OD) due to uncertain cognition caused by degraded visibility. With the strong assumption of data centralization, unsupervised domain adaptation (UDA) reduces vulnerability under OD scenario. Whereas, enlarged domain gap and growing privacy concern heavily challenge conventional UDA. Motivated by gap decomposition and data decentralization, we establish a decentralized domain adaptation (DDA) framework calledTranslate thEnAdapt (abbr.TEA) for privacy preservation. Our highlights lie in. (1) Regarding federated hallucination translation, aDisentanglement andContrastive-learning basedGenerativeAdversarialNetwork (abbr.DisCoGAN) is proposed to impose contrastive prior and disentangle latent space in cycle-consistent translation. To yield domain hallucination, client minimizes cross-entropy of local classifier but maximizes entropy of global model to train translator. (2) Regarding source-free regularization adaptation, aPrototypical-knowledge basedRegularizationAdaptation (abbr.ProRA) is presented to align joint distribution in output space. Soft adversarial learning relaxes binary label to rectify inter-domain discrepancy and inner-domain divergence. Structure clustering and entropy minimization drive intra-class features closer and inter-class features apart. Extensive experiments exhibit efficacy of our TEA which achieves 55.26% or 46.25% mIoU in adaptation from GTA5 to Foggy Cityscapes or Foggy Zurich, outperforming other DDA methods for SFSU. Xiating Jin, Jiajun Bu, Hui Zhang 0023, Yaonan Wang 0001 |
IEEE Trans. Multim. | 5 |
| 2025 | SSPD: Spatial-Spectral Prior Decoupling Model for Spectral Snapshot Compressive ImagingabstractCoded aperture snapshot spectral imaging (CASSI) captures 3D hyperspectral images (HSIs) in a single shot by encoding incident light into 2D measurements. However, recovering the original hyperspectral data from these measurements is a severely ill-posed inverse problem due to significant information loss during compression. Recent deep learning methods, especially deep unfolding networks, have demonstrated promising reconstruction results by embedding learnable priors into iterative optimization frameworks. However, most existing approaches use a single network to jointly estimate spatial and spectral priors, limiting their ability to handle the distinct properties of HSIs. To overcome this limitation, we propose the Spatial-Spectral Prior Decoupling Model (SSPD), which reformulates HSI reconstruction as a prior absorption problem, enabling independent modeling of spatial and spectral priors with specialized network architectures. To achieve this, we design two attention mechanisms tailored for hyperspectral data: one for capturing spatial correlations and another for preserving spectral signatures. Additionally, we develop a hybrid loss function that combines convergence constraints and cross-prior interactions, ensuring accurate prior fusion and stable reconstruction. Experiments on synthetic and real-world datasets confirm that SSPD outperforms existing methods in spectral snapshot compressive imaging. Lizhu Liu, Yaonan Wang 0001, Yurong Chen 0003, Jiwen Lu, Hui Zhang 0023 |
IEEE Trans. Multim. | 2 |
| 2025 | Cross-Modality Semantic Consistency Learning for Visible-Infrared Person Re-IdentificationabstractVisible-infrared person re-identification (VI-ReID) seeks to identify and match individuals across visible and infrared ranges within intelligent monitoring environments. Most current approaches predominantly explore a two-stream network structure that extract global or rigidly split part features and introduce an extra modality for image compensation to guide networks reducing the huge differences between the two modalities. However, these methods are sensitive to misalignment caused by pose/viewpoint variations and additional noises produced by extra modality generating. Within the confines of this articles, we clearly consider addresses above issues and propose a Cross-modality Semantic Consistency Learning (CSCL) network to excavate the semantic consistent features in different modalities by utilizing human semantic information. Specifically, a Parsing-aligned Attention Module (PAM) is introduced to filter out the irrelevant noises with channel-wise attention and dynamically highlight the semantic-aware representations across modalities in different stages of the network. Then, a Semantic-guided Part Alignment Module (SPAM) is introduced, aimed at efficiently producing a collection of semantic-aligned fine-grained features. This is achieved by incorporating parsing loss and division loss constraints, ultimately enhancing the overall person representation. Finally, an Identity-aware Center Mining (ICM) loss is presented to reduce the distribution between modality centers within classes, thereby further alleviating intra-class modality discrepancies. Extensive experiments indicate that CSCL outperforms the state-of-the-art methods on the SYSU-MM01 and RegDB datasets. Notably, the Rank-1/mAP accuracy on the SYSU-MM01 dataset can achieve 75.72%/72.08%. Min Liu 0008, Yuan Bian 0002, Yeqing Sun, Baida Zhang, Yaonan Wang 0001 |
IEEE Trans. Multim. | 7 |
| 2025 | Category-Level Multi-Object 9D State Tracking Using Object-Centric Multi-Scale Transformer in Point Cloud StreamabstractCategory-level object pose estimation and tracking has achieved impressive progress in computer vision, augmented reality, and robotics. Existing methods either estimate the object states from a single observation or only track the 6-DoF pose of a single object. In this paper, we focus on category-level multi-object 9-Dimensional (9D) state tracking from the point cloud stream. We propose a novel 9D state estimation network to estimate the 6-DoF pose and 3D size of each instance in the scene. It uses our devised multi-scale global attention and object-level local attention modules to obtain representative latent features to estimate the 9D state of each object in the current observation. We then integrate our network estimation into a Kalman filter to combine previous states with the current estimates and achieve multi-object 9D state tracking. Experiment results on two public datasets show that our method achieves state-of-the-art performance on both category-level multi-object state estimation and pose tracking tasks. Furthermore, we directly apply the pre-trained model of our method to our air-ground robot system with multiple moving objects. Experiments on our collected real-world dataset show our method's strong generalization ability and real-time pose tracking performance. Yaonan Wang 0001, Mingtao Feng, Huimin Lu 0002, Xieyuanli Chen |
IEEE Trans. Multim. | 2 |
| 2025 | Registration of Multiview Point Clouds With Unknown OverlapabstractRegistration of multiview point clouds obtained from 3D scanners is a common method for 3D reconstruction. However, most existing registration methods are designed to handle point clouds with known overlap relationships that are ensured by external equipment (e.g., manipulators, turntables) or acquisition sequences, which limits the application range and increases the acquisition cost. To overcome these limitations, an unknown overlap registration (UOR) method for multiview point clouds is proposed, which can estimate overlap confidence, construct a connected graph, and remove outlier point clouds automatically. First, the overlap confidence between two point clouds is estimated by calculating the average nearest neighbor feature distance within the predicted overlap region. We then construct a minimal spanning tree based on the confidence levels and search for the central node to serve as the world coordinate. Finally, the Lie algebra-based SE(3)-sensitive perturbation scheme is introduced to solve the fine transformations, in which a robust weighting function is designed to weight point correspondences. Our method can find reliable connections among point clouds, and the proposed graph can be combined with different pairwise registration methods. The experimental results on both indoor and industrial datasets demonstrate the accuracy and effectiveness of our method. Jiawen Zhao, Qing Zhu 0003, Yaonan Wang 0001, Weixing Peng, Hui Zhang 0023, Jianxu Mao |
IEEE Trans. Multim. | 3 |
| 2025 | Guided Adversarial Attack in the Low-Frequency SpaceabstractAdversarial examples can assess the robustness of machine learning models, which has attracted the attention of many researchers to adversarial example generation methods. Transferability and imperceptibility stand out as two crucial metrics for evaluating the quality of adversarial examples. However, achieving a balance between these two indicators poses a formidable challenge. In this paper, we propose a low-frequency guided adversarial attack method (LGA) to generate adversarial examples with strong transferability and good imperceptibility. Specifically, we enhance the transferability of adversarial examples by increasing the diversity of attack algorithms, and introduce the guiding principle and the triplet loss constraint to ensure that the generated adversarial examples are optimized away from the class regions of the clean examples. We find that the low-frequency component in the frequency domain of the image contains the vast majority of the semantic information of the image. Therefore, we constrain the attack perturbations to low-frequency component space to enhance the covert nature while maintaining visual coherence, rendering the adversarial examples more difficult to perceive. We conduct extensive experiments on various models with different network structures and multiple defense strategies, and the experimental results demonstrate that our method outperforms existing methods in the tradeoff between transferability and imperceptibility, achieving the SOTA performance. Lingping Tan, Yanchun Li, Shujuan Tian, Yaonan Wang 0001 |
IEEE Trans. Multim. | 6 |
| 2025 | Spatially Covariant Image Registration With Text PromptsabstractMedical images are often characterized by their structured anatomical representations and spatially inhomogeneous contrasts. Leveraging anatomical priors in neural networks can greatly enhance their utility in resource-constrained clinical settings. Prior research has harnessed such information for image segmentation, yet progress in deformable image registration has been modest. Our work introduces textSCF, a novel method that integrates spatially covariant filters and textual anatomical prompts encoded by visual-language models, to fill this gap. This approach optimizes an implicit function that correlates text embeddings of anatomical regions to filter weights. textSCF not only boosts computational efficiency but can also retain or improve registration accuracy. By capturing the contextual interplay between anatomical regions, it offers impressive interregional transferability and the ability to preserve structural discontinuities during registration. textSCF's performance has been rigorously tested on intersubject brain magnetic resonance imaging (MRI) and abdominal computerized tomography (CT) registration tasks, outperforming existing state-of-the-art models in the MICCAI Learn2Reg 2021 challenge and leading the leaderboard. In abdominal registrations, textSCF's larger model variant improved the Dice score by 11.3% over the second-best model, while its smaller variant maintained similar accuracy but with an 89.13% reduction in network parameters and a 98.34% decrease in computational operations. Xiang Chen 0008, Min Liu 0008, Rongguang Wang, Renjiu Hu, Gaolei Li, Yaonan Wang 0001, Hang Zhang 0010 |
IEEE Trans. Neural Networks Learn. Syst. | 7 |
| 2025 | Contrastive Learning Framework With Cross-Sensor Adaptive Signal Representation for Fault DiagnosisabstractAlthough multisource sensor (MS) signal-based mechanical fault diagnosis (MFD) can significantly improve the diagnostic performance, the existing methods often lack sufficient adaptability and generalization when retraining on single-sensor signals or inferring from partial sensor signals. Thus, a general two-stage signal representation contrastive learning fault diagnosis framework (T-SCF) is proposed to adapt the trained model to varying numbers of sensor signals. This framework enhances model robustness and data fusion by comparing sensor signal views, offering a new approach for information fusion, fault detection, and classification in MFD. In the first stage, an adaptive contrastive algorithm is proposed to generate contrastive samples (C-Ss) and contrastive labels (C-Ls) for MS signals. Then, a supervised contrastive loss (SCL) is designed to minimize the similarity between different fault MS signals while maximizing the similarity between identical ones. By designing a parallel encoder architecture, SCL enables it to merge contrasting the features of different sensor signals during training. This strategy preserves the time-domain dimension properties of different sensors during the training of the second-stage classifier, thereby improving the adaptability of the model to different sensor signals without affecting the global information. The effectiveness of the method was verified from multiple different evaluation dimensions using two public datasets and one self-built dataset. Jianxu Mao, Yaonan Wang 0001, Zhe Li 0050, Hui Zhang 0023 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2025 | Pixel-Level Noise Mining for Weakly Supervised Salient Object DetectionabstractTraining a deep model for visual saliency detection requires the collection and labor-intensive annotation of overwhelmingly large data. We propose to learn saliency detection in a weakly supervised manner from single noisy label, which is easy to obtain from unsupervised handcrafted feature-based methods. However, deep networks tend to overfit such noises leading to a dramatic drop in accuracy. Given our goal, we address a natural question: can we identify outliers during network prediction and rectify the label noises? To this end, we propose a pixel-level noise mining framework for robust salient object detection (SOD) by exploiting its own knowledge, and without the need for external models. Specifically, during the early training stage, we progressively identify the outliers from a novel perspective during saliency detection, before the network overfits to the noisy labels, and generate a selection matrix in each iteration. Next, we adaptively rectify the label noises under the guidance of the selection matrix for better supervision in the later training stage. Extensive experiments on multiple benchmark datasets demonstrate the superiority of our method showing its ability to learn saliency detection comparable to state-of-the-art fully supervised methods. Furthermore, our approach outperforms existing weakly supervised methods utilizing single noisy label and surpasses the half of existing weakly supervised methods employing multiple noisy labels. Our approach, which trains with multiple noisy labels, outperforms all other methods employing multiple noisy labels across four major datasets. Furthermore, we also evaluate the generalization ability of our method on the multiclass semantic segmentation (SS) task. Our code is available at https://github.com/kendongdong/NoiseMining. Kendong Liu, Mingtao Feng, Wei Zhao 0019, Weisheng Dong, Yaonan Wang 0001, Ajmal Mian |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2025 | MBUNeXt: Multibranch Encoder Aggregation Network Based on Layer-Fusion Strategy for Multimodal Brain Tumor SegmentationabstractMultimodal brain tumor segmentation (BraTS), integrated with surgical robots and navigation systems, enables accurate surgical interventions while maximizing the preservation of surrounding healthy brain tissue. However, multimodal brain scans suffer from large interclass differences in brain tumor subregions and information redundancy, leading to inadequate fusion of multimodal information and significantly affecting the accuracy of BraTS. To address the above problems, we propose a multibranch encoder aggregation (MEA) network based on a layer-fusion strategy called multibranch UNeXt (MBUNeXt). The network comprises three well-designed modules: the multimodal feature attention (MFA) module, the MEA module, and the large-kernel convolution skip (LCS)-connection module. These modules work together to achieve precise segmentation of brain tumors. Specifically, the MFA module preserves the intermodality similarity structure through attention mechanisms and Gaussian modulation functions, thereby filtering redundant information. Then, the MEA module exploits the correlations among multiple modalities to effectively integrate multimodal hybrid feature representation and optimize multimodal information fusion. In addition, the LCS module constructs multiple groups of depthwise separable convolutions with large kernel, which can guide the network to attend to features at different scales, thereby addressing the issue of significant interclass differences in brain tumor subregions. The experimental results on the large-scale public datasets, BraTS2019 and BraTS2021, which consist of approximately 5000 3-D brain scans, demonstrate that our proposed method has achieved SOTA performance, with average Dice scores of 85.84% and 91.11%, respectively. It also performs well on the BraTS-Africa2024 dataset with low imaging quality, confirming its robustness. The code is available at https://github.com/liuqinghao2018/MBUNeXt. Qinghao Liu, Yuehao Zhu, Min Liu 0008, Zhao Yao, Yaonan Wang 0001, Erik Meijering |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2025 | Grouped Vector Autoregression Reservoir Computing Based on Randomly Distributed Embedding for Multistep-Ahead PredictionabstractAs an efficient recurrent neural network (RNN), reservoir computing (RC) has achieved various applications in time-series forecasting. Nevertheless, a poorly explained phenomenon remains as to why the RC and deep RCs succeed in handling time-series prediction despite completely randomized weights. This study tries to generate a grouped vector autoregressive RC (GVARC) time-series forecasting model based on the randomly distributed embedding (RDE) theory. In RDE-GVARC, the deep structures are constructed by multiple GVARCs, which makes the established RDE-GVARC evolve into a deterministic deep RC model with few hyperparameters. Then, the spatial output information of the GVARC is mapped into the future temporal states of an output variable based on RDE equations. The main advantages of the RDE-GVARC can be summarized as follows: 1) RDE-GVARC solves the problems of uncertainty in the weight matrix and difficulty in large-scale parameter selection in the input and hidden layers of deep RCs; 2) the GVARC can avoid massive deep RC hyperparameter design and make the design of deep RC more straightforward and effective; and 3) the proposed RDE-GVARC shows good performance, strong stability, and robustness in several chaotic and real-world sequences for multistep-ahead prediction. The simulating results confirm that the RDE-GVARC not only outperforms some recently deep RCs and RNNs, but also maintains the rapidity of RC with an interpretable structure. Heshan Wang, Zhepeng Wang 0003, Mingyuan Yu, Jing J. Liang, Jinzhu Peng, Yaonan Wang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2025 | Learning Granularity-Aware Affordances From Human-Object Interaction for Tool-Based Functional Dexterous GraspingabstractTo enable robots to use tools, the initial step is teaching robots to employ dexterous gestures for touching specific areas precisely where tasks are performed. Affordance features of objects serve as a bridge in the functional interaction between agents and objects. However, leveraging these affordance cues to help robots achieve functional tool grasping remains unresolved. To address this, we propose a granularity-aware affordance feature extraction method for locating functional affordance areas and predicting dexterous coarse gestures. We study the intrinsic mechanisms of human tool use. On the one hand, we use fine-grained affordance features of object-functional finger contact areas to locate functional affordance regions. On the other hand, we use highly activated coarse-grained affordance features in hand-object interaction regions to predict grasp gestures. Additionally, we introduce a model-based postprocessing module that transforms affordance localization and gesture prediction into executable robotic actions. This forms GAAF-Dex, a complete framework that learns granularity-aware affordances from human-object interaction to enable tool-based functional grasping with dexterous hands. Unlike fully supervised methods that require extensive data annotation, we employ a weakly supervised approach to extract relevant cues from exocentric (Exo) images of hand-object interactions to supervise feature extraction in egocentric (Ego) images. To support this approach, we have constructed a small-scale dataset, functional affordance hand (FAH)-object interaction dataset, which includes nearly 6k images of functional hand-object interaction Exo images and Ego images of 18 commonly used tools performing six tasks. Extensive experiments on the dataset demonstrate that our method outperforms state-of-the-art methods, and real-world localization and grasping experiments validate the practical applicability of our approach. The source code and the established dataset are available at https://github.com/yangfan293/GAAF-DEX. Fan Yang 0063, Wenrui Chen, Kailun Yang 0001, Conghui Tang, Zhiyong Li 0001, Yaonan Wang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 8 |
| 2025 | Distributed Neural Adaptive Impedance Control for Cooperative Manipulation With Unknown ObjectsabstractExisting cooperative manipulation methods for multiple manipulator systems usually assume that the grasp matrix and the desired trajectory of each manipulator are known in advance. In this work, distributed neural adaptive impedance control (AIC) strategies integrating fully distributed observers are proposed to remove both limitations. Specifically, two fully distributed finite-time observers are designed to estimate the actual and ideal states of the reference point without using global information. The estimates of the grasp matrix and the desired trajectory of each end-effector (EE) are then obtained by kinematic constraints and the estimates of the reference point's states. At the controller development, a distributed adaptive impedance model is established to achieve an adaptive trade-off between tracking performance and compliance. Then, distributed neural network (NN)-based tracking control strategies are developed to asymptotically realize the desired adaptive impedance dynamics in the presence of uncertainties. Additionally, a virtual energy tank (EK) is employed to interact with the impedance system to correct the adaptive impedance laws for system passivity. A simulation for four mobile manipulators tightly cooperative transport an unknown object is carried out to demonstrate the established results. Danping Zeng, Yaonan Wang 0001, Yiming Jiang 0001, Haoran Tan, Zhiqiang Miao, Yun Feng 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2025 | Reliable Wind Turbine Blade Performance Monitoring System Using Aerodynamic Audio Signals and Deep Learning ApproachesabstractWind turbines have emerged as a prominent and environmentally friendly energy generation solution. However, with the widespread use of new materials, ensuring the reliability of these devices has become as a critical issue. Developing efficient and cost-effective monitoring methods for the wind turbine's blades (WTBs), the most expensive components of wind turbine, has become a focal point of research. In this article, we present a novel monitoring system for WTBs that employs a deep convolutional neural network approach based on the medical auscultatory method. The system is designed to balance economic efficiency and engineering reliability. First, we proposed a lightweight WTBs monitoring framework based on edge computing that leverages the signals from the programmable logic controller output of wind turbine to enable efficient collection of relevant aerodynamic audio signals while filtering out irrelevant data. Second, we present a set of audio enhancement algorithms that employ multiscale feature extraction, self-adaptive mask targeting, and deep neural networks to reduce noise in the audio signals generated by WTBs. Third, we introduce a new approach for compressing deep convolution neural networks that makes them suitable for resource-constrained edge computing devices and efficiently utilizes audio-generated spectrograms to diagnose faults in WTBs. Baheti Biekezat, Hui Zhang 0023, Yihong Cao, Yurong Chen 0003, Yaonan Wang 0001 |
IEEE Trans. Reliab. | 5 |
| 2025 | A Local Knowledge Transfer-Based Evolutionary Algorithm for Constrained Multitask OptimizationabstractEvolutionary multitask optimization (EMTO) can solve multiple tasks simultaneously by leveraging the relevant information between tasks, but existing EMTO algorithms do not take into account the fact that almost all problems in the real world contain constraints. To address this dilemma, this article studies a local knowledge transfer-based evolutionary algorithm for constrained multitask optimization. To be specific, each task population is divided into multiple niches to enhance the diversity and control the intensity of knowledge transfer, thus avoiding excessive transfer of knowledge. Then a new similarity judgment method based on the information feedback of pioneer individuals is developed to judge the similarity between tasks and whether to perform knowledge transfer. Furthermore, two different transfer methods: a direct transfer and a learning transfer, are devised to perform knowledge transfer among niches pertaining to different tasks. In addition, an excellent-information-guided mutation mechanism is proposed to prevent niches from getting trapped in local optima and to promote rapid convergence. The system experiment on 18 constrained multitask test instances and 2 real-world problems demonstrate that the proposed algorithm outperforms or is at least comparable to other EMTO algorithms and constrained single-objective optimization algorithms. Xuanxuan Ban, Jing J. Liang, Kunjie Yu, Yaonan Wang 0001, Kangjia Qiao, Jinzhu Peng, Dun-Wei Gong, Canyun Dai |
IEEE Trans. Syst. Man Cybern. Syst. | 4 |
| 2025 | CIMAP: A High-Performance Motion Planning Algorithm for Robotic Manipulators in Complex Environments Using Clearance Inference NetworkabstractThis article introduces CIMAP, a high-performance motion planning algorithm for robotic manipulators in complex environments, based on the clearance inference network (CIN). CIMAP incorporates a batch collision estimation module powered by CIN, which efficiently predicts collisions by dividing the manipulator’s workspace into voxels and estimating clearances between the manipulator and surrounding obstacles. The algorithm also features a batch adaptive bidirectional expansion mechanism, enabling the simultaneous extension of multiple nodes within joint space. Leveraging CIN for batch collision estimation, CIMAP accelerates the discovery of feasible paths. Additionally, CIMAP includes a phased path optimization mechanism that identifies local shortcuts through CIN, improving path efficiency. A geometric collision checker ensures safety, performing necessary repairs when required. To assess CIMAP’s effectiveness in continuous motion planning, we compared its performance against four existing algorithms (CN-RRT, B-RRT, GB-RRT*, and NPB-RRT*-DC) across various obstacle scenarios. Experimental results demonstrate that CIMAP achieves an average motion planning time of under 0.7 s, improving planning efficiency by at least 89% compared to the baseline algorithms, while maintaining shorter path lengths. Bo Chen 0047, Hui Zhang 0023, Fangfang Zhang 0004, Yiming Jiang 0001, Wei He 0001, Chenguang Yang 0001, Yaonan Wang 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 7 |
| 2025 | A Predefined-Time Robust Sliding Mode Control Based on Zeroing Neural Dynamics for Position and Attitude Tracking of QuadrotorabstractSliding mode control (SMC) is considered an efficacious scheme for quadrotor control. However, the control performance of the existing SMC schemes depends on initial states and multiple parameters, and the robustness needs to be improved. To address these issues, a novel predefined-time robust SMC framework based on two zeroing neural dynamics (ZND) schemes, referred to as ZND-based predefined-time robust SMC (ZNDPRSMC) framework, is developed to facilitate position and attitude tracking of a quadrotor under bounded disturbances. Initially, a nonsingular sliding mode surface (SMS) is formulated by incorporating a general ZND along with a differentiable predefined-time activation function. Following this, an approaching law is introduced by utilizing a variable-parameter noise-tolerant ZND and a novel dynamic adaptive parameter. The nonsingular SMS and the approaching law are then combined to construct a nonsingular predefined-time robust controller. The theoretical proofs provided ascertain the predefined-time convergence of the closed-loop system utilizing ZNDPRSMC and its robustness against bounded disturbances. Finally, two trajectory tracking examples of the quadrotor are presented to demonstrate the superiority of the ZNDPRSMC framework. Yongjun He 0001, Lin Xiao 0002, Qiuyue Zuo, Hang Cai, Yaonan Wang 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 5 |
| 2025 | A Novel Hierarchical Distributed Robust Formation Control Strategy for Multiple Quadrotor AircraftsabstractMultiple quadrotor aircraft system has significant advantages for performing complex tasks in dangerous environments, but it is still challenging for formation with external disturbance or internal model uncertainty. This article establishes a hierarchical distributed robust formation strategy for multiple quadrotor aircrafts, in which trajectory tracking and attitude formation are, respectively, controlled in different layers. An upper trajectory tracking controller is proposed to generate desired position for lower anti-disturbance attitude formation controller, where a bi-level adaptive terminal sliding mode controller with disturbance observers are, respectively, designed. In lower control layer, the desired velocity is generated by velocity control part to ensure quadrotor aircraft to track desired position while maintaining certain formation shape, whereas the acceleration control part is responsible for driving actual velocity of each quadrotor aircraft to desired velocity. Stability analysis shows that the prescribed formation can be realized if unknown disturbance is bounded and time constants in different layers are selected appropriately. Compared with existing results, the proposed strategy enables multiple complex tasks to be realized in different layers to improve formation accuracy and achieve interference suppression. The effectiveness is verified through both Gazebo simulation and actual experiment. Qianxiong Li, Xiaoqing Lu, Yaonan Wang 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 3 |
| 2025 | A Similar-Niching-Based Differential Evolution for Constrained Multimodal Multiobjective OptimizationabstractIn constrained multimodal multiobjective optimization problems (CMMOPs), the existence of discrete and confined feasible regions bring great challenges to current multiobjective optimization evolutionary algorithms (MOEAs). To address these challenges, this article proposes a constrained multimodal multiobjective differential evolution algorithm, which incorporates a similar-niching-based reproduction operator and a novel environmental selection mechanism. The proposed algorithm initiates by segregating the population into distinct niches, thereby promoting independent evolution within each niche. This segmentation enhances the exploration of multiple discrete feasible regions, thus improving the capacity to find diverse Pareto optimal solutions. Moreover, the algorithm selects the most similar niche to collaboratively generate solutions, further enhancing its ability to generate effective feasible solutions. To improve the diversity within the population, the proposed environmental selection mechanism gives preference to solutions that enhance the distribution of the next-generation population. By considering the diversity in both two spaces, the population retains more pareto optimal solutions. Based on the Friedman test results of the comparison experiment with other representative algorithms and the champion algorithm of the CEC2023 CMMOPs competition, the proposed algorithm attained the top ranking, thereby reinforcing its demonstrated superiority. Meanwhile, the proposed algorithm is used to solve the constrained multimodal multiobjective location selection problem and results show its superiority. Jing J. Liang, Caitong Yue, Ying Bi 0001, Kangjia Qiao, Yaonan Wang 0001, Ponnuthurai N. Suganthan |
IEEE Trans. Syst. Man Cybern. Syst. | 6 |
| 2025 | Deep spatial and discriminative feature enhancement network for stereo matching
Guowei An, Yaonan Wang 0001, Kai Zeng 0010, Qing Zhu 0003, Xiaofang Yuan |
Vis. Comput. | 2 |
| 2024 | A Niching-Based Reproduction and Preselection-Based Multiobjective Differential Evolution for Multimodal Multiobjective OptimizationabstractIn multimodal multiobjective optimization problems (MMOPs), there are several Pareto optimal solutions corresponding to the identical objective vector. MMOPs pose greater challenges for multiobjective optimization evolutionary algorithms (MOEAs) as they require balancing the convergence and diversity of the population in both the decision space and the objective space. Therefore, this paper proposes a novel coevolutionary multiobjective optimization differential evolution algorithm with a niching-based reproduction and a preselection-based environmental selection mechanism, called NPCMODE. The algorithm introduces a niching-based reproduction, evolving solutions independently within multiple niches to generate more dispersed solutions in the decision space. Additionally, a preselection-based environmental selection mechanism priori-tizes solutions with low density in both decision and objective spaces through a dual-population coevolutionary framework. The efficiency of NPCMODE is validated through comparative experiments with six representative multimodal multiobjective optimization evolutionary algorithms (MMOEAs) on the CEC 2019 benchmark suite, showcasing its effectiveness in achieving a balance between convergence and diversity performance. Jing J. Liang, Caitong Yue, Yaonan Wang 0001 |
CEC | 4 |
| 2024 | L4D-Track: Language-to-4D Modeling Towards 6-DoF Tracking and Shape Reconstruction in 3D Point Cloud Streamabstract3D visual language multi-modal modeling plays an important role in actual human-computer interaction. However, the inaccessibility of large-scale 3D-language pairs restricts their applicability in real-world scenarios. In this paper, we aim to handle a real-time multi-task for 6-DoF pose tracking of unknown objects, leveraging 3D-language pre-training scheme from a series of 3D point cloud video streams, while simultaneously performing 3D shape reconstruction in current observation. To this end, we present a generic Language-to-4D modeling paradigm termed L4D-Track, that tackles zero-shot 6-DoF Tracking and shape reconstruction by learning pairwise implicit 3D representation and multi-level multi-modal alignment. Our method constitutes two core parts. 1) Pairwise Implicit 3D Space Representation, that establishes spatial-temporal to language coherence descriptions across continuous 3D point cloud video. 2) Language-to-4D Association and Contrastive Alignment, enables multi-modality semantic connections between 3D point cloud video and language. Our method trained exclusively on public NOCS-REAL275 dataset, achieves promising results on both two publicly benchmarks. This not only shows powerful generalization performance, but also proves its remarkable capability in zero-shot inference. The project is released at L4D- Track. Yaonan Wang 0001, Mingtao Feng, Yulan Guo, Ajmal Mian, Zheng Shou 0001 |
CVPR | 2 |
| 2024 | External Knowledge Enhanced 3D Scene Generation from Sketch
Mingtao Feng, Yaonan Wang 0001, He Xie, Weisheng Dong, Bo Miao, Ajmal Mian |
ECCV (6) | 3 |
| 2024 | Safety-Critical Control for Underwater Vehicles with Model Uncertainties and External DisturbancesabstractSafety is a crucial issue for underwater vehicles, which may be affected by narrow terrain and multiple obstacles. In addition, the model of the underwater vehicles are uncertain and susceptible to external disturbances such as water flow. This article utilizes model predictive control (MPC) and incremental nonlinear dynamic inversion (INDI) to design a robust control scheme for underwater vehicles. The position loop controller employs MPC to generate the required speed commands for the velocity loop controller. The velocity loop is designed with an INDI control scheme incorporating a second-order low-pass filter, effectively mitigating model uncertainties and external disturbances on the vehicles. Based on exponential control barrier functions (ECBFs), the input constraint and obstacle avoidance problems of underwater vehicles are solved. The results indicate that the proposed control scheme not only exhibits robustness but also effectively ensures safe obstacle avoidance. Yizong Chen, Zhiqiang Miao, Weiwei Zhan, Yaonan Wang 0001 |
ICARCV | 4 |
| 2024 | NMPC for Trajectory Tracking of Hybird Terrestrial-Aerial Vehicles with Collision AvoidanceabstractIn recent years, researches on Hybrid Terrestrial-Aerial Vehicles (HTAVs) have received a lot of attention. However, most of the existing researches focus on achieving basic motion control in both modes, thus neglecting the situation of encountering obstacles during motion. To fill this research gap and achieve accurate trajectory tracking with collision avoidance, we proposes the use of Nonlinear Model Predictive Control (NMPC). In this paper, we firstly focus on passive-wheeled HTAVs as the target platform. Subsequently, we introduce distance constraint function and dynamics models for both aerial and terrestrial modes, along with three motion constraints specific to the terrestrial mode. Then we designed the NMPC controller base on distance constraint function and the dynamics models for each mode, incorporating the motion constraint in the terrestrial controller to ensure smooth movement on the ground. At the end of this paper, we conduct simulation experiments to evaluate the effectiveness of trajectory tracking. The results of simulations demonstrate that regardless of the mode, the HTAV achieves a relatively high tracking accuracy and avoid collision when following a nonlinear trajectory with fixed initial position and orientation. Additionally, the position error converges rapidly and exhibits minimal fluctuation, highlighting the significant role played by the added constraints in control. Zhiqiang Miao, Haoming Tang, Yizong Chen, Yaonan Wang 0001 |
ICARCV | 5 |
| 2024 | Distributed Resilient Estimator for Networked Systems Under Deception AttacksabstractDeception attacks are employed to compromise cyber-physical systems through fake data injection. This paper concentrates on the distributed resilient estimation issue of multi-sensor networked systems under deception attacks. In order to detect deception attacks, we utilize Kullback-Leibler(K-L) divergence as a criterion to distinguish the discrepancy between the deceived information and the estimated information. When the attack does not exist, the transmitted information can be restored to ensure the resilient estimation performance. Based on the extended Kalman filter design method, a distributed resilient estimation with a dual-gain mechanism is developed. This advanced approach dynamically adjusts the weighting balance between the predictive model and sensor data inputs, achieving the optimal estimation during the shutdown and activation of spoofing attacks. Finally, numerical simulations are provided to further illustrate the results. Weiwei Zhan, Zhiqiang Miao, Yizong Chen, Yaonan Wang 0001 |
ICARCV | 6 |
| 2024 | Domain Adaptation in Visual Reinforcement Learning via Self-Expert Imitation with Purifying Latent FeatureabstractGeneralizing visual reinforcement learning is fundamental to robot visual navigation, involving the acquisition of a policy from interactions with source environments to facilitate adaptation to analogous, yet unfamiliar target environments. Recent advancements capitalize on data augmentation techniques, self-supervised learning methods, and the generative adversarial network framework to train policy neural networks with enhanced generalizability. However, current methods, upon extracting domain-general latent features, further utilize these features to train the reinforcement learning policy, resulting in a decline in the performance of the learned policy guiding the agent to accomplish tasks. To tackle these challenges, a framework of self-expert imitation with purifying latent features was devised, empowering the policy to achieve robust and stable zero-shot generalization performance in visually similar domains previously unseen, without diminishing the performance of guiding the agent to accomplish tasks. The extraction method of domain-general latent features is proposed to enhance their quality based on the variational autoencoder. Extensive experiments have shown that our policy, compared with state-of-the-art counterparts, does not diminish the performance of the policy guiding the agent to accomplish tasks after generalization. Lin Chen 0034, Jianan Huang 0002, Zhen Zhou 0003, Yaonan Wang 0001, Yang Mo, Zhiqiang Miao, Kai Zeng 0010, Mingtao Feng, Danwei Wang |
IROS | 4 |
| 2024 | Decentralized Multi-Robot Navigation Coupled with Spatial-Temporal RetNet Based on Deep Reinforcement LearningabstractNavigating robots through dynamic multi-robot environments, avoiding collisions with both other robots and obstacles, has emerged as a central challenge in robotics. The existing approaches fall short in allowing the policy network to effectively capture spatial-temporal reciprocal collision avoidance in multi-robot environments, comprising both static and dynamic obstacles, resulting in inadequate safety and efficiency in directing robot movement. In this study, we introduce a novel policy neural network called Spatial-Temporal RetNet (STR), designed to encode reciprocal collision avoidance states between robots in spatial and temporal dimensions. The goal is to improve the safety and efficacy of the policy neural network in directing robots to complete assigned tasks. The spatial state encoder module is built upon a parallel RetNet structure, which strengthens the neural network's capacity in extracting reciprocal collision avoidance states between robots in spatial dimensions. This module addresses the limitations of position encoding in transformer-based multi-robot navigation policy neural networks. We design a temporal state encoder utilizing a recurrent RetNet structure. This innovation bolsters the multi-robot navigation policy neural network's capability to capture features in the temporal dimension of multi-robot movements. It addresses the limitations of transformer-based multi-robot navigation policy neural networks, particularly in recurrently inferring information across time dimensions. Simulation experiments were conducted to showcase the superior safety and effectiveness of our proposed method compared to previous state-of-the-art approaches in guiding robots to accomplish tasks. Lin Chen 0034, Yaonan Wang 0001, Zhiqiang Miao, Mingtao Feng, Yuanzhe Wang, Yang Mo, Zhen Zhou 0003, Hesheng Wang 0001, Danwei Wang |
IROS | 2 |
| 2024 | Decentralized Trajectory Planning for Formation Flight in Unknown and Dense EnvironmentsabstractFor aerial swarms, formation flight has been applied in various scenes. However, most existing works do not consider balancing the conflicting requirements among keeping formation, keeping the smoothness of trajectories, and obstacle avoidance within the limited time. To address this issue, we propose a decentralized trajectory planning framework for formation flight in unknown and dense environments. To ensure that feasible trajectories can be found within the limited time, the formation optimization problem is decoupled into formation affine transformation and iterative trajectory generation. Firstly, the optimization problem based on affine transformation is designed to obtain the optimal affine transformation sequence, which provides the formation reference of trajectory optimization. Secondly, the iterative optimization framework of trajectory planning is designed, which balances the conflicting requirements of formation, smooth flight, and obstacle avoidance. Besides, to escape the local minima caused by non-convex dense environments, the method of topological path planning is designed to provide distinctive initial solutions for trajectory optimization. Finally, the proposed methods are proven to be effective through the simulations and real-world experiments. Jianxin Zeng, Yaonan Wang 0001, Zhiqiang Miao, Wei He 0001, Hesheng Wang 0001 |
IROS | 2 |
| 2024 | Interpretable Unsupervised Homography Estimation
Zhen Zhou 0003, Qing Zhu 0003, Yaonan Wang 0001, Yang Mo, Lin Chen 0034, Jianan Huang 0002, Tianjian Jiang |
PRCV (2) | 3 |
| 2024 | MHA-DGCLN: multi-head attention-driven dynamic graph convolutional lightweight network for multi-label image classification of kitchen waste
Qiaokang Liang, Hai Qin, Mingfeng Liu, Dongbo Zhang 0003, Yaonan Wang 0001, Dan Zhang 0006 |
Appl. Intell. | 7 |
| 2024 | Discriminative target predictor based on temporal-scene attention context enhancement and candidate matching mechanism
Baiheng Cao, Xuedong Wu, Xianfeng Zhang, Yaonan Wang 0001 |
Expert Syst. Appl. | 4 |
| 2024 | A transformer-based lightweight method for multiple-object trackingabstractAbstract At present, the multi‐object tracking method based on transformer generally uses its powerful self‐attention mechanism and global modelling ability to improve the accuracy of object tracking. However, most existing methods excessively rely on hardware devices, leading to an inconsistency between accuracy and speed in practical applications. Therefore, a lightweight transformer joint position awareness algorithm is proposed to solve the above problems. Firstly, a joint attention module to enhance the ShuffleNet V2 network is proposed. This module comprises the spatio‐temporal pyramid module and the convolutional block attention module. The spatio‐temporal pyramid module fuses multi‐scale features to capture information on different spatial and temporal scales. The convolutional block attention module aggregates channel and spatial dimension information to enhance the representation ability of the model. Then, a position encoding generator module and a dynamic template update strategy are proposed to solve the occlusion. Group convolution is adopted in the input sequence through position encoding generator module, with each convolution group responsible for handling the relative positional relationships of a specific range. In order to improve the reliability of the template, dynamic template update strategy is used to update the template at the appropriate time. The effectiveness of the approach is validated on the MOT16, MOT17, and MOT20 datasets. Qin Wan 0001, Zhu Ge, Yang Yang 0052, Xuejun Shen, Hang Zhong, Hui Zhang 0023, Yaonan Wang 0001, Di Wu 0046 |
IET Image Process. | 7 |
| 2024 | A video object detector with Spatio-Temporal Attention Module for micro UAV detection
Haozhi Xu, Zhigang Ling, Xiaofang Yuan, Yaonan Wang 0001 |
Neurocomputing | 4 |
| 2024 | Asymptotic stability analysis of time delayed fractional-order replicator dynamics with government's intervention
Zhang Zhe, Toshimitsu Ushio, Yaonan Wang 0001, Jing Zhang 0014, Xiaogang Zhang 0002 |
Neurocomputing | 3 |
| 2024 | Autonomous obstacle avoidance and target tracking of UAV: Transformer for observation sequence in reinforcement learning
Weilai Jiang, Tianqing Cai, Yaonan Wang 0001 |
Knowl. Based Syst. | 4 |
| 2024 | Structural and Textural-Aware Feature Extraction for Hyperspectral Image ClassificationabstractFeature extraction is a prevalent technique in hyperspectral remote sensing. Various tasks require this technique as a pre-processing step, including image classification, anomaly detection, image denoising, and so on. Edge-preserving filtering based methods have been extensively utilized for this purpose. However, these methods do not take the inherent structural and textural information into account, leading to poor performance in classifying hyperspectral images (HSIs). In this letter, a new structural and textural-aware feature extraction method is proposed that preserves the relevant structural information and removes useless textures. First, structural and textural-aware recursive filtering features (STRFs) are extracted along with an exponential form of windowed inherent variance (eWIV). Then, multi-scale STRFs are integrated by the principal component analysis (PCA) method to obtain more discriminative features (MSTRF). Finally, the fused features are fed into a pixel-wise classifier to obtain the final results. The main difference between the MSTRF method and other feature extraction methods is that the MSTRF method can make full use of the proposed eWIV map, which can help to properly characterize structure and texture in HSIs. Experimental results on several public data sets indicate that our method leads to state-of-the-art classification performance, especially in the presence of very small training set. Ying Zhang 0063, Lianhui Liang, Jun Li 0009, Antonio Plaza, Xudong Kang, Jianxu Mao, Yaonan Wang 0001 |
IEEE Geosci. Remote. Sens. Lett. | 7 |
| 2024 | Visual tracking via confidence template updating spatial-temporal regularized correlation filters
Mengquan Liang, Xuedong Wu, Siming Tang, Yaonan Wang 0001, Baiheng Cao |
Multim. Tools Appl. | 5 |
| 2024 | A Two-Stage Noise-Tolerant Paradigm for Label Corrupted Person Re-IdentificationabstractSupervised person re-identification (Re-ID) approaches are sensitive to label corrupted data, which is inevitable and generally ignored in the field of person Re-ID. In this paper, we propose a two-stage noise-tolerant paradigm (TSNT) for labeling corrupted person Re-ID. Specifically, at stage one, we present a self-refining strategy to separately train each network in TSNT by concentrating more on pure samples. These pure samples are progressively refurbished via mining the consistency between annotations and predictions. To enhance the tolerance of TSNT to noisy labels, at stage two, we employ a co-training strategy to collaboratively supervise the learning of the two networks. Concretely, a rectified cross-entropy loss is proposed to learn the mutual information from the peer network by assigning large weights to the refurbished reliable samples. Moreover, a noise-robust triplet loss is formulated for further improving the robustness of TSNT by increasing inter-class distances and reducing intra-class distances in the label-corrupted dataset, where a constraint condition for reliability discrimination is carefully designed to select reliable triplets. Extensive experiments demonstrate the superiority of TSNT, for instance, on the Market1501 dataset, our paradigm achieves 90.3% rank-1 accuracy (6.2% improvement over the state-of-the-art method) under noise ratio 20%. Min Liu 0008, Fei Wang 0124, Yaonan Wang 0001, Amit K. Roy-Chowdhury |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2024 | Weakly Supervised Tracklet Association Learning With Video Labels for Person Re-IdentificationabstractSupervised person re-identification (re-id) methods require expensive manual labeling costs. Although unsupervised re-id methods can reduce the requirement of the labeled datasets, the performance of these methods is lower than the supervised alternatives. Recently, some weakly supervised learning-based person re-id methods have been proposed, which is a balance between supervised and unsupervised learning. Nevertheless, most of these models require another auxiliary fully supervised datasets or ignore the interference of noisy tracklets. To address this problem, in this work, we formulate a weakly supervised tracklet association learning (WS-TAL) model only leveraging the video labels. Specifically, we first propose an intra-bag tracklet discrimination learning (ITDL) term. It can capture the associations between person identities and images by assigning pseudo labels to each person image in a bag. And then, the discriminative feature for each person is learned by utilizing the obtained associations after filtering the noisy tracklets. Based on that, a cross-bag tracklet association learning (CTAL) term is presented to explore the potential tracklet associations between bags by mining reliable positive tracklet pairs and hard negative pairs. Finally, these two complementary terms are jointly optimized to train our re-id model. Extensive experiments on the weakly labeled datasets demonstrate that WS-TAL achieves 88.1% and 90.3% rank-1 accuracy on the MARS and DukeMTMC-VideoReID datasets respectively. The performance of our model surpasses the state-of-the-art weakly supervised models by a large margin, even outperforms some fully supervised re-id models. Min Liu 0008, Yuan Bian 0002, Qing Liu 0035, Yaonan Wang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2024 | Cross-lingual font style transfer with full-domain convolutional attention
Tian-le Ji, Paul L. Rosin, Yukun Lai, Weiliang Meng, Yaonan Wang 0001 |
Pattern Recognit. | 6 |
| 2024 | HairManip: High quality hair manipulation via hair element disentangling
Lin Zhang 0041, Paul L. Rosin, Yukun Lai, Yaonan Wang 0001 |
Pattern Recognit. | 5 |
| 2024 | Toward Safe Distributed Multi-Robot Navigation Coupled With Variational Bayesian ModelabstractDesigning a safe and effective collision avoidance policy for multiple robots is essential in decentralized scenarios, where each robot is responsible for generating its own paths, to ensure their safe operation. Recently, the utilization of reinforcement learning to develop decentralized policies that enable multiple robots to move cooperatively and accomplish tasks has yielded positive outcomes. However, the presence of exploration unsafe actions during the reinforcement learning training process results in inadequate safety. We seek to enhance the safety of distributed multi-robot navigation policies and propose a new imitation learning framework based on the variational Bayesian model, which enables robots to learn safe actions by anticipating the subsequent state they are expected to reach. In addition, a new policy neural network structure for multi-robot navigation is proposed by introducing the transformer structure, which encodes the significance of nearby robots in relation to their forthcoming conditions. Experiments demonstrated that our policy can more safely guide robots to navigate in multi-robot environments under conditions of limited information, outperforming the state-of-the-art RL-RVO method in terms of success rate.Note to Practitioners—The motivation of this paper is to address the problem of collision avoidance in a multi-robot environment under limited information, which can also be applied to autonomous driving, crowd simulation, and other related fields. Positive outcomes have been observed in the utilization of reinforcement learning to create decentralized policies that enable multiple robots to move cooperatively and complete tasks. However, inadequate safety remains a challenging task due to the possibility of exploring hazardous actions during training. This article aims to enhance the safety of distributed policies guiding robots to accomplish navigation tasks in dynamic multi-robot environments. To begin with, we introduce a novel framework for imitation learning that is based on the variational Bayesian model. This framework facilitates the learning of safe actions by the policy to improve its performance and guide the robot in navigating and avoiding obstacles more securely. A loss function is proposed that enables the anticipation of the future state expected to be reached by the robot. By incorporating the transformer structure, a new neural network structure is designed for multi-robot navigation that encodes the significance of nearby robots concerning their upcoming conditions. This network structure employs a BiGRUs to facilitate the assimilation of observations from multiple agents by the policy. Compared to existing works such as GA3C-CADRL, SARL, and RL-RVO, our proposed method achieves a higher success rate. In our future research, we will investigate methods to enhance the policy’s performance in guiding robots to complete tasks by focusing on improving travel time and average speed, while also strictly ensuring safe navigation. Furthermore, we plan to extend this approach by addressing navigation challenges in more densely populated multi-robot environments. Lin Chen 0034, Yaonan Wang 0001, Zhiqiang Miao, Mingtao Feng, Zhen Zhou 0003, Hesheng Wang 0001, Danwei Wang |
IEEE Trans Autom. Sci. Eng. | 2 |
| 2024 | Multiple Mobile Robots Planning Framework for Herding Non-Cooperative TargetabstractNon-cooperative target herding is one of the concerns in the robotics field. For the non-cooperative target herding problem, a general planning framework is proposed to compel the target to the destination by a mobile robot team with pursuit, encirclement, and guidance operations. In the proposed planning framework, the encirclement strategy and guidance strategy are designed to deal with the unpredictability of the target. Firstly, the mobile robots approach the randomly moving target in the pursuit process. Secondly, the encirclement strategy is applied in the encirclement process to enable the mobile robot team to encircle the non-cooperative target. Finally, the guidance strategy is employed in the guidance process to enable the mobile robot team to form a favorable encirclement formation and compel the non-cooperative target to the destination. During the herding task, the mobile robots encircle the target or other mobile robots in a satellite-like motion. Moreover, two queues (orbiters and wanders) are maintained to change the formation adaptively, improving the flexibility of the framework. Various environments are designed to verify the effectiveness of the planning framework both in simulations and real-world experiments. Furthermore, the proposed framework is a general platform having great potential for various applications, where the numbers of mobile robots in the herding team are changeable, and different planning algorithms can be integrated into the framework.Note to Practitioners—The motivation for this paper is to propose a planning framework to efficiently herd the non-cooperative target in various scenarios for applications such as the maintenance of public safety and the facilitation of wildlife migration. The existing cooperative herding methods typically maintain relatively fixed formations, while the planning framework proposed in this paper introduces a satellite-based motion pattern and diverse formation transformation strategies, thereby enhancing formation flexibility and adaptability. Consequently, this framework can be effectively applied to challenging environments such as those with restricted areas and dynamic obstacles. Furthermore, the proposed planning framework allows for flexible adjustments in path planning methods and the number of mobile robots to meet various practical application requirements. Through a series of simulations and real-world experiments, the reliability and favorable performance of the proposed algorithm framework in practical applications are demonstrated. Yangning Wu, Bingwei He, Hui Zhang 0023, Yaonan Wang 0001 |
IEEE Trans Autom. Sci. Eng. | 7 |
| 2024 | Adaptive Force Tracking Impedance Control for Aerial Interaction in Uncertain Contact Environment Using Barrier FunctionabstractIn this article, an adaptive force tracking impedance control strategy is investigated for an aerial manipulator in physical interaction with uncertain contact environments. Based on the modified target impedance model, an adaptive impedance control method is proposed to accomplish aerial interaction in uncertain environments while maintaining a stable contact force, wherein the environment parameters of location and stiffness are estimated online to generate a reference position trajectory. Then, in order to ensure the tracking performance of the aerial manipulator, a robust pose tracking controller is designed, including a barrier function-based position controller and an adaptive attitude controller. Both proposed position and attitude controllers can ensure finite-time convergence of the state variable without the priori boundary information of disturbances. In particular, the position state variable can converge to a predefined neighborhood of zero from any initial state, and the control gain is not overestimated. The stability of the proposed strategy is analyzed via Lyapunov tools. Simulations and real-world experiments are conducted to illustrate the feasibility and performance of the proposed control strategy.Note to Practitioners—The motivation of this article is to investigate an adaptive force tracking impedance control strategy for aerial physical interaction with uncertain contact environments. In the existing impedance control schemes for aerial manipulators, the environment parameter of location or stiffness is often required to be utilized in controller design. However, in practical cases, the environmental parameters are not known precisely. Thus, this article presents an adaptive impedance method to automatically generate the reference position trajectory and achieve a stable contact force. Additionally, the tracking performance of the aerial manipulator is inevitably subject to uncertainties and disturbances. To ensure tracking convergence, traditional robust controllers generally involve high control gains than the known upper bounds of the disturbances. The main disadvantage of those controllers is that the control gain is often overestimated when the disturbance decreases. To address this issue, a barrier function-based position controller is proposed for the aerial manipulator, where the priori boundary information of disturbances is not needed and the control gain is adaptively adjusted according to the amplitude of disturbances. The stability and convergence of the proposed strategy are analyzed mathematically, and the experiments using an aerial manipulator provide promising results. Jiacheng Liang, Hang Zhong, Yaonan Wang 0001, Junhao Zeng, Jianxu Mao |
IEEE Trans Autom. Sci. Eng. | 3 |
| 2024 | Robust High-Order Control Barrier Functions-Based Optimal Control for Constrained Nonlinear Systems With Safety-Stability PerspectivesabstractIn this article, we propose a robust high-order control barrier functions (HoCBFs)-based optimal control method for nonlinear systems with state constraints to achieve safety-stability perspectives. First, a kind of HoCBFs is presented for constrained nonlinear systems to address state constraints with high relative degrees. Second, the robustness property of the HoCBFs is analyzed based on the asymptotic stability of the forward invariant set. Specifically, a robust HoCBFs-based Lyapunov function is constructed to prove the uniform asymptotic stability of the set associated with the HoCBFs. In this way, a new sufficient condition is obtained for the stability analysis of the forward invariant set by using the inequalities of high-order derivatives of Lyapunov function. Third, a robust HoCBFs-based optimal control scheme is proposed for the constrained nonlinear system to achieve the safety-stability perspectives of constraints satisfaction and system stabilization, where the robust HoCBFs are combined with control Lyapunov functions (CLFs) to satisfy the small control property (SCP) in solving a quadratic program (QP). Furthermore, the proposed optimal control scheme is shown to be Lipschitz continuous and has no initial condition restrictions. Finally, two examples are presented to demonstrate the control performance of the proposed scheme.Note to Practitioners—The motivation of this article is that constraints exist widely in actual control systems, and the lack of constraint satisfaction in control systems may inevitably lead to safety defects, which usually degrade the control performances or even damage the entire system. In this article, a robust HoCBFs-based optimal control scheme is proposed for constrained nonlinear systems. The theoretical derivation demonstrates that the proposed control scheme can achieve safety-stability perspectives, which ensure system stabilization and task-oriented performance without violating the state constraints. The satisfactory control performances of the simulation on a constrained robotic manipulator show the potential practical application on a real robotic system. Jinzhu Peng, Haijing Wang, Shuai Ding 0007, Jing J. Liang, Yaonan Wang 0001 |
IEEE Trans Autom. Sci. Eng. | 5 |
| 2024 | Viewpoint Planning of Robotic Measurement System for Free-Form Surfaces Based on Visibility Cone Space ExplorerabstractFree-form surfaces have been widely used in industrial design and manufacturing. For the requirements of measurement efficiency and precision, robots and optical scanners are applied to measure free-form surface parts increasingly. Due to the complex geometry shapes and occlusions of these parts, how to plan accessible viewpoints of a scanner to achieve the expected coverage rate is a challenging task. This paper presents a novel viewpoint planning method based on the visibility cone space explorer (VP-VCSE) for robotic measurement systems with 7 degrees of freedom (7-DOF). A digital twin for the robotic measurement system is implemented to provide core services for robotic measurement tasks, including sensor simulation and collision detection. To generate initial candidate viewpoints, a novel mesh segmentation algorithm based on the hybrid mixture model is proposed, which is convenient to handle the triangular mesh of the target object. Visibility computation for a target object in given viewpoints is the key to dealing with the occlusion problem. For this purpose, a general visibility model of a structured-light scanner is presented to compute visible areas accurately. In order to reduce occlusions, a visibility cone space explorer is designed to search optimal candidate viewpoints considering inverse kinematics and physical collisions simultaneously. The viewpoint planning problem is formulated as a set covering optimization problem and a next-best-view operator is introduced to improve the efficiency of the genetic algorithm for searching the resultant viewpoint set, guaranteeing the expected coverage rate and data overlap rate. The simulation and experiment results for four different test models show that the proposed algorithm outperforms the existing methods in terms of the uncovered rate and the minimum number of viewpoints.Note to Practitioners—This paper addressed a viewpoint planning problem for the robotic measurement system with a binocular structured light 3D scanner mounted on the end effector of the robot, where a robot and a turntable cooperate to complete the measurement tasks. The goal is to find a minimal number of viewpoints that provides full coverage of the target surfaces. Although many studies have addressed this problem, there is little discussion about strategies to improve coverage rate when the target object has complex occlusions. This paper suggested a valuable practice to construct a visibility cone space to adjust viewpoint to reduce occlusions and improve the overall coverage rate. Simulation and experimental results demonstrated the feasibility and effectiveness of the proposed approach. This paper showed how to deal with various constraints that a feasible viewpoint needs to satisfy in the viewpoint generation, viewpoint adjustment, and viewpoint selection phase. Moreover, this paper provided a solution for developing the visualization, simulation, and interaction of a digital twin for the 7-DOF robotic measurement system. All core services for robotic measurement tasks are implemented based on a set of open source libraries, which provides a convenient learning and research software platform for practitioners. In future research, we will study how to improve the intelligence and cooperation of the robotic measurement system through deep learning or reinforcement learning techniques. Yongpeng Tang, Yaonan Wang 0001, Haoran Tan, He Xie, Yiming Jiang 0001, Weixing Peng |
IEEE Trans Autom. Sci. Eng. | 2 |
| 2024 | Hybrid Force/Position Control for Switchable Unmanned Aerial Manipulator Between Free Flight and Contact OperationabstractThe refined aerial operations of the unmanned aerial manipulator (UAM) have been extensively studied for the last decades. Usually, UAM operations are accompanied by several phases, such as free flight, contact operations, and separation. A great challenge is proposed for switching operations in different environments and the high-precision contact force requirements of UAM control, so this paper conducts control stability research in the case of dynamic differences between free flight and contact operation for the UAM system. First, a hybrid force/position control strategy is proposed for switchable UAM system, among them, the adaptive sliding mode control method based on the interference observer is utilized in the free flight phase, and the adaptive impedance force control method is employed in the contact operation phase, where the adaptive estimation method is designed to perform on-board manipulator contact force estimation. Then a robust adaptive control strategy is proposed for the attitude loop to compensate for the torque disturbance generated during the contact operation phase. Meanwhile, the stability of the switching system is analyzed through the continuous Lyapunov function to prove the stability of the switching process. Finally, the effectiveness and superiority of the proposed schemes are verified through contact operation simulations and experiments.Note to Practitioners—This work is motivated by the contact force tracking of an UAM without force sensor. In recent years, hybrid force/position control has been widely used in UAM. However, force measurement is required at the end-effector and the object in most studies. The proposed method divides the contact operation process into free flight and contact operation stages. In the free flight stage, it is not necessary to know the prior information of the environment accurately (i.e. the external disturbance of slowly varying or known upper bound), and the disturbance observer is adopted to compensate for the disturbance caused by the external environment. In the contact operation stage, the impedance force control method is used for force tracking, which reduces the complexity and quality of the end-effector. At the same time, The shock caused by the switch from free flight to contact operation is reduced by calculating the appropriate controller parameters using the Lyapunov function. The proposed method is a promising solution for real applications and is validated via simulation and indoor contact experiments. The experimental results show that the proposed method has better stability and accuracy than the existing methods, and can be extended to industrial inspection, component processing, aerial operation, etc. Yangning Wu, Bingwei He, Zhiqiang Miao, Hui Zhang 0023, Yaonan Wang 0001 |
IEEE Trans Autom. Sci. Eng. | 7 |
| 2024 | Hybrid Force/Position Control of Multi-Mobile Manipulators for Cooperative Operation Without Force MeasurementsabstractIn this paper, considering the difficulty of the interaction between the multi-mobile manipulators and the environment, the dynamics model of the mobile manipulator is analyzed, and a hybrid force/position control method based on the prescribed performance is proposed to improve the stability of the multi-mobile manipulators in the process of cooperative object transportation. Firstly, the dynamics model of the underdriven system of the multi-mobile manipulators is established by the Newton-Euler theorem. Then, the equivalent control theory is adopted for underdriven system, and a prescribed performance control method is proposed by considering the motion interference between the mobile manipulator and the high precision control of the manipulator. At the same time, an adaptive impedance control method is used to overcome internal and external disturbance during the cooperative transport of multi-mobile manipulators. The stability of the proposed method is analyzed through the Lyapunov stability theory. Finally, the effectiveness and superiority of the proposed scheme are verified through a simulation of multi-mobile manipulators collaborative object transportation. Jianxu Mao, Haoran Tan, Yiming Jiang 0001, Yun Feng 0001, You Wu 0005, Yaonan Wang 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 7 |
| 2024 | Optimized Weights for Heterogeneous Epidemic Spreading Networks: A Constrained Cooperative Coevolution StrategyabstractThe wide spreading of COVID-19 all over the world raises numerous focus on the epidemic containment problem. Different from the traditional epidemic control strategies that focus on quarantine and vaccination, we seek to control the epidemic from a network science perspective, i.e., by adjusting the weights of the epidemic spreading networks. Moreover, considering the limitations on the available resources, the dynamic constrained optimization problem of weights’ adaptation for heterogeneous epidemic spreading networks is investigated. Due to the powerful ability of searching for global optimum, evolutionary algorithms (EAs) are used as optimizers. One major difficulty is that the dimension of the problem is increasing exponentially with the network size and most existing EAs cannot achieve satisfiable performance on large-scale optimization problems. To address this issue, a novel constrained cooperative coevolution ($C^3$) strategy, which can separate the original large-scale problem into different subcomponents, is used to achieve the tradeoff between the constraint and objective function. Yun Feng 0001, Yaonan Wang 0001, Bing-Chuan Wang, Li Ding 0013 |
IEEE Trans. Comput. Soc. Syst. | 2 |
| 2024 | Prior Images Guided Generative Autoencoder Model for Dual-Camera Compressive Spectral ImagingabstractCompressive Spectral Imaging (CSI) techniques have attracted considerable attention among researchers for their ability to simultaneously capture spatial and spectral information using low-cost, compact optical components. A prominent example of CSI techniques is the Dual-Camera Coded Aperture Snapshot Spectral Imaging (DC-CASSI), which involves reconstructing hyperspectral images from CASSI measurements and uncoded panchromatic or RGB images. Despite its significance, the reconstruction process in DC-CASSI is challenging. Conventional DC-CASSI techniques rely on different models to explore the similarity between uncoded images and hyperspectral images. Nevertheless, two main issues persist: i) the effective utilization ofspatial informationfrom RGB images to guide the reconstruction process, and ii) the enhancement ofspectral consistencyof recovered images when using panchromatic/RGB images, which inherently lack precise spectral information. To address these challenges, we propose a novel Prior images guided generative autoEncoder (PiE) model. The PiE model leverages RGB images as prior information to enhance spatial details and designs a generative model to improve spectral quality. Notably, the generative model is optimized in a self-supervised manner. Comprehensive experimental results demonstrate that the proposed PiE method outperforms existing techniques, achieving state-of-the-art performance. Yurong Chen 0003, Yaonan Wang 0001, Hui Zhang 0023 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | 3D Object Detection From Point Cloud via Voting Step Diffusionabstract3D object detection is a fundamental task in scene understanding. Numerous research efforts have been dedicated to better incorporate Hough voting into the 3D object detection pipeline. However, due to the noisy, cluttered, and partial nature of real 3D scans, existing voting-based methods often receive votes from the partial surfaces of individual objects together with severe noises, leading to sub-optimal detection performance. In this work, we focus on the distributional properties of point clouds and formulate the voting process as generating new points in the high-density region of the distribution of object centers. To achieve this, we propose a new method to move random 3D points toward the high-density region of the distribution by estimating the score function of the distribution with a noise conditioned score network. Specifically, we first generate a set of object center proposals to coarsely identify the high-density region of the object center distribution. To estimate the score function, we perturb the generated object center proposals by adding normalized Gaussian noise, and then jointly estimate the score function of all perturbed distributions. Finally, we generate new votes by moving random 3D points to the high-density region of the object center distribution according to the estimated score function. Extensive experiments on two large scale indoor 3D scene datasets, SUN RGB-D and ScanNet V2, demonstrate the superiority of our proposed method. The code will be released athttps://github.com/HHrEtvP/DiffVote. Haoran Hou, Mingtao Feng, Weisheng Dong, Qing Zhu 0003, Yaonan Wang 0001, Ajmal Mian |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2024 | Category-Contextual Relation Encoding Network for Few-Shot Object DetectionabstractFew-shot object detection (FSOD) has brought increasing academic interest by recognizing previously unseen novel classes with very limited well-labeled samples. However, most existing methods identify novel classes via some object-specific characteristics in the few provided samples rather than intrinsic inter-class relations between base and novel classes, which heavily degrades the detection performance on novel classes. Moreover, they cannot learn discriminative proposal representations to distinguish base and novel classes, and thus misclassify novel objects as confusable base classes. To tackle the above challenges, we develop a novel Category-contextual Relation Encoding Network (CRE-Net), which is an early attempt to reason inter-class context relationships for FSOD task. To be specific, we propose a novel category-contextual relation encoding mechanism to capture intrinsic inter-class relations between base and novel classes via knowledge aggregation from global category-contextual descriptors. It utilizes intrinsic inter-class contextual relations to adaptively refine the convolution kernel, thus encoding the local semantic context of query image with category-contextual relation as guidance. Furthermore, to explore discriminative representations for base and novel classes, we develop a scarcity-compensatory contrastive proposal loss by incorporating data scarcity of novel classes and proposal semantic consistency with high confidence. This loss could compact object instances from the same category to a tighter cluster, and enhance the space separability of different classes. Extensive experiments on Pascal VOC and COCO datasets verify the state-of-the-art detection performance of our CRE-Net model when compared with other baseline methods. Ating Yin, Yaonan Wang 0001, Jianxu Mao, Hui Zhang 0023, Xiuyi Chen |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | Deep Stereo Network With MRF-Based Cost AggregationabstractDespite the remarkable progress made in learning-based stereo-matching algorithms, it is an open challenge for stereo-matching in disparity discontinuities and textureless regions. In this paper, we propose the deep Markov Random Field based cost aggregation network (DMCA-Net) for stereo matching, which is an end-to-end model-driven network architecture. This architecture introduces an efficient feature extraction network to extract richer textual and contextual feature information for stereo feature similarity representation at multi-stages and levels. Furthermore, with the aim of alleviating the edge-fattening phenomenon at disparity discontinuities and generating accurate disparities in textureless regions, we proposed the differentiable Markov Random Field model for cost aggregation, where the model’s data term utilizes image detail information, such as boundary and contour features, to guide matching cost aggregation, and the model’s smoothness term penalizes the adjacency similarity of the cost between the four-nearest neighboring pixel pairs to predict the disparity in textureless regions. The detailed experiment demonstrates that DMCA network achieves competitive performance on the SceneFlow, KITTI 2012, KITTI 2015, and Middlebury 2014 datasets. Kai Zeng 0010, Hui Zhang 0023, Wei Wang 0025, Yaonan Wang 0001, Jianxu Mao |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | Unsupervised Homography Estimation With Pixel-Level SVDDabstractHomography estimation is a common image alignment method. Unsupervised learning, which uses unlabeled training and exhibits excellent performance, has attracted much attention in this field. When there are multiple planes in the scene, using features over the entire image for matching will lead to compromised results. However, existing methods for learning focused principal plane masks through deep neural networks lack explicit guidance. In this paper, we propose a novel unsupervised method to explicitly model anomaly descriptor removal and mask generation. Specifically, reliable feature descriptors are selected from a novel perspective, and regard the features that are not responsible for alignment as outliers. The pixel-level support vector data description (PL-SVDD) module is designed. This module learns the feature representation of image pixels and fits a hypersphere to exclude the feature redundancy information that is not responsible for alignment from the hypersphere, thereby optimizing the feature descriptor. Based on the optimized image features, a correlation learning (CL) module is designed. This module displays a generated mask through mathematical modeling to select reliable areas for homography estimation. Specifically, the feature descriptor of one unaligned images is modeled as a multivariate Gaussian distribution by Gaussian density estimation (GDE). Then, The Mahalanobis distance is combined with the multivariate Gaussian distribution of the model and the feature descriptor of another image to generate the mask. Experiments show that our method achieves good performance compared with previous methods. Zhen Zhou 0003, Qing Zhu 0003, Mingtao Feng, Yaonan Wang 0001, Jianqiao Luo, Zhiqiang Miao, Lin Chen 0034, Yang Mo |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | A Novel Zeroing Neurodynamic Method Based on Discrete Fuzzy Control System: Design, Analysis, and VerificationabstractConsidering the extensive research on zeroing neurodynamic (ZN), a self-adaptive and enhanced fixed-time convergent zeroing neurodynamic (SEFC-ZN) method for addressing time-variant problems is presented in this paper based on a discrete fuzzy matrix (DFM) design parameter and a novel advanced sign-bi-power activation function (NASbpAf). Due to the distinctive design of the DFM design parameter and NASbpAf, the proposed SEFC-ZN method possesses prominent self-adaptivity and enhanced fixed-time convergence. Specifically, the DFM design parameter is actually a matrix with all elements generated from a discrete fuzzy control system, so it can self-adaptively adjust the convergence rate of every error in the SEFC-ZN method resulting in the self-adaptivity. This feature is greatly different from the conventional scalar design parameters whose values are usually fixed or increase indefinitely and different errors in the ZN method can only be adjusted by the same design parameter. By summarizing the characteristic of the activation functions designed previously according to the SbpAf, it is found that keeping two terms of the SbpAf and adding extra terms can improve the performance of the ZN method. Thereout, built on the SbpAf, the NASbpAf is presented which can make the SEFC-ZN method realize the enhanced fixed-time convergence. Three theoretical analyses and proofs, together with relative corollaries, conclude the properties of the SEFC-ZN method and the advantages of the DFM design parameter and NASbpAf. A numerical experiment about solving time-variant nonlinear equations by the SEFC-ZN method and an application to the linear-quadratic optimal control strongly verify the proposed theory and method. Lei Jia 0001, Lin Xiao 0002, Yaonan Wang 0001, Jianhua Dai 0003, Biao Luo 0001 |
IEEE Trans. Fuzzy Syst. | 3 |
| 2024 | Robust Image-Based Adaptive Fuzzy Controller for Guarantee Field of View With Uncertain DynamicsabstractVisual servoing technology has widely been employed in manufacturing because it is a flexible, realizability, and low-cost way to improve the intelligence of the industry robot. Nevertheless, a worrisome and overlooked issue is that the loss of visual features in the camera's field of view may lead to the failures of the visual servoing tasks. This article addresses the visual features escaping problem, by implementing an asymmetric barrier Lyapunov function with a field-of-view constraint controller. The asymmetric barrier Lyapunov function defines a tightly specified range for the feature coordinate errors and ensures the transient response of the tracking error as well as enables arbitrary tracking accuracy. It is worth noting that the asymmetric barrier Lyapunov function directly handles the visual-robot-coupled dynamics while guaranteeing system stabilities. Besides, to accommodate the uncertain dynamics derived from a high-dimensional coupled system, an adaptive controller is proposed utilizing fuzzy neural networks with computational efficiency and few training parameters to enhance the control performance. Finally, the effectiveness of the proposed control strategy has been demonstrated through both theoretical analysis and experimental verification. Jiao Jiang, Yaonan Wang 0001, Yiming Jiang 0001, Yun Feng 0001, Hang Zhong, Chenguang Yang 0001 |
IEEE Trans. Fuzzy Syst. | 2 |
| 2024 | Neural-Network-Based Security Control for T-S Fuzzy System With Cooperative Event-Triggered MechanismabstractIn this article, we investigate the problem of security control for T-S fuzzy Markov jump systems (FMJSs) under actuator faults and deception attacks, while introducing a cooperative event-triggered mechanism (CETM). In order to enhance the efficiency of communication resources, we develop the CETM in the forward channel, which operates concurrently on sensor-to-observer (STO) and observer-to-controller (OTC) channels utilizing a united event generator. Additionally, we design an attack-compensating controller to eliminate the impact of nonlinear malicious injection information generated by deceptive attacks on the system, where the compensation signal is generated by approximating the attack signal using radial basis function neural network (RBFNN) technology. Furthermore, using the Lyapunov function, sufficient conditions for ensuring that T-S FMJSs are mean square exponential ultimate bounded (MSEUB) are derived. Finally, the effectiveness of our proposed approach is demonstrated through a simulation example. Cheng Tan 0001, Chengzhen Gao, Jinzhu Peng, Xiangpeng Xie 0001, Yaonan Wang 0001 |
IEEE Trans. Fuzzy Syst. | 5 |
| 2024 | Learning-Based Resilient Adaptive Fuzzy Optimal Consensus for Nonlinear Multiagent Systems Under DoS AttacksabstractThis study addresses the learning-based resilient adaptive fuzzy optimal consensus control problem for nonlinear uncertain Multiagent Systems (MASs) in the presence of intermittent Denial of Service (DoS) attacks. A key obstacle is the uncertainty in the dynamics of the followers, which makes it challenging to eliminate dependency on the identifier network. To this end, we propose a novel critic-only optimal consensus scheme to eliminate dependency on the identifier network and significantly reduce computational complexity. Moreover, this work requires less prior knowledge and assumes that only the specific subsystems can access the leader's information under certain conditions. To cope with limited information access, we design a distributed adaptive observer to monitor the leader's dynamics. It is proven that all the signals are uniformly ultimately bounded(UUB), and consensus tracking is achieved. Finally, a simulation example is provided to demonstrate the results achieved. Meijian Tan, Zhi Liu 0001, Yaonan Wang 0001, C. L. Philip Chen, Zongze Wu 0001 |
IEEE Trans. Fuzzy Syst. | 3 |
| 2024 | EGST: An Efficient Solution for Human Gaits Recognition Using Neuromorphic Vision SensorabstractTraditional cameras struggle to perform in challenging scenarios such as low latency, high speed and high dynamic range. In contrast, neuromorphic vision sensors (event cameras) have great potential for robotics and computer vision due to the advantages of high temporal resolution, high dynamic range, and ultra-low resource consumption. Event cameras are novel bio-inspired sensors that monitor the brightness change of each pixel asynchronously and provide a stream of events encoding the time, position and sign of the brightness changes. Hence, traditional computer vision methods cannot be directly applied to the event-stream. Finding event representations that completely maintain event attributes, as well as efficient and accurate learning approaches, is the key to unlocking the potential of event cameras. In this study, we reveal the rigid transfer from event-stream to graph that has been overlooked in previous work and introduce a novel event representation, namely event graph sequence (EGS) considering the local and global temporal clues. Coupled with EGS, we propose a spatio-temporal pattern extracting (STPE) module to capture the spatio-temporal correlation and evolution of EGS. Our novel framework, Event Graph Sequence Transformer (EGST), exploits event properties to provide efficient and accurate recognition. This study focuses on the event-based human gaits recognition task, and EGST is evaluated on three different event-based gait datasets. The evaluation results show better or comparable accuracy than the state-of-the-art, while requiring extremely low computation resources. The code will be available athttps://github.com/C19h/EGST. Liaogehao Chen, Zhenjun Zhang, Yang Xiao 0007, Yaonan Wang 0001 |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2024 | ADMM-DSP: A Deep Spectral Image Prior for Snapshot Spectral Image DemosaicingabstractSpectral imaging, with the ability to simultaneously capture the spectral and spatial information of scenes, has obtained researchers' significant interest. Traditional spectral imaging techniques typically suffer from high costs and slow imaging speed. Multispectral filter array-based snapshot imaging is a cutting-edge technology for mitigating these problems. The spectral image demosaicing algorithm plays a key role in reconstructing high-resolution spectral images from the raw measurement. In this article, we first formulate the multispectral demosaicing problem as a compressive spectral imaging problem. Then, a novel deep spectral image prior is introduced as the regularization, which assumes that neural networks can generate the desired spectral image from the raw image in a self-supervised learning manner. Finally, the constructed constrained problem is solved by the alternating direction method of multipliers optimization algorithm. Compared with existing methods, experiments on various scenes demonstrate that the proposed method obtains superior performance. Yurong Chen 0003, Hui Zhang 0023, Yaonan Wang 0001, Ating Yin |
IEEE Trans. Ind. Informatics | 3 |
| 2024 | A Novel Methodology to Predict 3-D Surface Temperature Field on Delamination for ThermographyabstractThis article proposes a novel 3-D surface temperature prediction model based on the restored pseudoheat flux (RPHF) theory. The method can be used to simulate the temperature difference between the subsurface defect and the sound area. The proposed model shows the potential to investigate the detection limits associated with the defect features, such as depth, radius, diameter-to-depth ratio (D2dR), and excitation features, which is beneficial for the experimental design. Several experiments were conducted on specimens of different materials [glass fiber reinforced plastic (GFRP), CFRP, and rubber] using RPHF thermography to validate the practicality of the model. The comparative analysis is also conducted with other methods. Both experimental and simulation results demonstrated that longer heating is required for deeper defects and the moment of maximum temperature difference tends to appear after the heating has stopped. Probability of detection (PoD) was used as an index to assess the reliability of the methodology and the problems found. The depth of the defect has a greater influence on thermal detection than D2dR. Xiang Li 0159, Hongjin Wang, Yunze He, Yaonan Wang 0001 |
IEEE Trans. Ind. Informatics | 4 |
| 2024 | Robust Variable Impedance Control for Aerial Compliant Interaction With Stability GuaranteeabstractThis article investigates a robust variable impedance control methodology for aerial manipulators to realize compliant and safe interaction tasks. Considering that the stability characteristics are generally overlooked in existing variable impedance controllers of the aerial manipulator, state-independent stability conditions are applied for time-varying impedance profiles to ensure the exponential stability of the desired variable impedance dynamics (DVID) as well as the boundedness of the state variables in the DVID. A command trajectory variable is introduced for converting the impedance control issue to a particular tracking issue, and then, a robust variable impedance controller based on the wrench estimator is designed to guarantee the exponential convergence of the translational states and impedance error of the aerial manipulator. The designed impedance controller is structurally simple and results in low implementation costs. Next, an improved attitude control approach with the command filter is developed for global flight attitude stability without any singularities or ambiguities, where the filter is introduced to avoid computing the derivative signals of the generalized force input. Finally, the effectiveness of the proposed control method is illustrated via numerical simulations and interaction experiments with different targets in real scenarios. Jiacheng Liang, Yaonan Wang 0001, Hang Zhong, Hongwen Li, Jianxu Mao, Wei Wang 0025 |
IEEE Trans. Ind. Informatics | 2 |
| 2024 | A Depth Adaptive Feature Extraction and Dense Prediction Network for 6-D Pose Estimation in Robotic GraspingabstractEstimating the 6-D pose of an object is a vital and challenging task for robot vision systems in industrial robotic grasping. With the wide use of 3-D cameras, the additional acquired depth image provides geometric information of the scene to increase the pose estimation performance but leads to a challenge, fully leveraging the two-modal data, the color image and the depth image. Previous works usually adopt two individual strategies to handle the data, which suffer from limited accuracy and efficiency since the two complementary data are not fully explored. Thus, we propose a depth adaptive feature extraction and dense prediction network that decouples the scale-dependent and the scale-invariant information from the depth image. The former guides the network to perceive the 3-D structure of the scene, and the latter, together with color image, provides the scene textures for feature extraction. The proposed network not only fuses multimodal textures but also retains their 3-D structure. In addition, a dense prediction strategy is adopted to regress the object pose; this approach can mitigate the instability caused by outliers. We conduct various evaluations on a real-world industrial dataset to illustrate the advantages of the proposed approach; and a practical robotic grasping platform is presented to demonstrate its application performance. Xuebing Liu, Xiaofang Yuan, Qing Zhu 0003, Yaonan Wang 0001, Mingtao Feng, Zhen Zhou 0003 |
IEEE Trans. Ind. Informatics | 4 |
| 2024 | MSRN: Multilevel Spatial Refinement Network for Transmission Line Fastener Defect DetectionabstractTransmission line (TL) fasteners play the role of connecting components in smart-grid transmission processes with abnormal TL fastener states, seriously impacting the power supply. Therefore, regular detection of TL fasteners is significant. However, the images taken by unmanned aerial vehicle have problems, such as small size and complex background, which bring great challenges to the existing object detection models. Based on this, this article proposes a multilevel spatial refinement network (MSRN), including an attention-guided receptive field enhanced feature pyramid network (ARFE-FPN) and a double refinement head (DR-Head). For the small target problem, ARFE-FPN first uses dilated convolution to expand the receptive field, and uses global average pooling to extract background activation values. Then, it performs channel weighting on TL fastener features under different fields of view. For the problem of complex background, DR-Head first constructs a semantic prediction task to realize the preseparation of foreground and background, and then combines the high-resolution feature map to further highlight the fastener features in the low-resolution feature map. Experiments on the TL fastener dataset show that MSRN has the best detection accuracy, and its AP can reach 92$\%$. Jianxu Mao, Qingxian Liu, Yaonan Wang 0001, Weixing Peng, Junfei Yi, Ziming Tao, Hui Zhang 0023, Caiping Liu |
IEEE Trans. Ind. Informatics | 3 |
| 2024 | Safe Obstacle Avoidance Planning-Control Scheme for Multiconstrained Mobile ManipulatorsabstractTo achieve high precision and safety in the operation of a wheeled mobile manipulator, it is imperative that the robot possesses the capability for high-precision tracking while adhering to multiple physical constraints and avoiding obstacles. This article introduces a novel approach that combines model predictive control (MPC) with prescribed performance function (PPF) to address these challenges. At the kinematic level, we leverage MPC's predictive capabilities to optimize the robot's motion for a better reference velocity while taking into account the preestablished velocity tracking error bounds defined by PPF. On the dynamics level, the control law is designed based on PPF, ensuring precise tracking of the reference velocity and desired end-effector trajectory. The validity and effectiveness of this method are rigorously validated through a series of empirical experiments. Jingmou Nie, Yaonan Wang 0001 |
IEEE Trans. Ind. Informatics | 2 |
| 2024 | Deep Correspondence Matching-Based Robust Point Cloud Registration of Profiled PartsabstractDue to ability to estimate the spatial transformation of coordinate frames, point cloud registration is a fundamental technique in manufacturing. Previous methods prone to converge to wrong local minima, in the cases of large initialization, noise, outliers, and partiality. This article presents a new learning-based robust point cloud registration approach to predict a rigid transformation in a one-shot way. Our network aims to determine a matchability matrix to yield an accurate registration result. Each element of the matchability matrix refers to similarity of learned per-point embeddings and represents the probability of a potential correspondence. The following two major blocks are developed to guide the matchability matrix to represent correct correspondences: an attention block is introduced to enhance the discriminativeness of learned per-point embeddings, and a zero-mean Gaussian-based annealing layer and a differentiable Sinkhorn normalization layer are designed to enforce a permutation matchability matrix. With the matchability matrix, an intuitive solution is integrated to obtain the relative transformation of the source and target point clouds. Different from the existing work, our network can handle partially overlapped point-cloud pairs effectively. Experimental results demonstrate the superiority of the proposed approach over the state-of-the-art registration approaches in terms of accuracy and robustness. Weixing Peng, Yaonan Wang 0001, Hui Zhang 0023, Yihong Cao, Jiawen Zhao, Yiming Jiang 0001 |
IEEE Trans. Ind. Informatics | 2 |
| 2024 | A Fast Tracking Network for Pedestrian Following of Mobile Robot in Unknown Complex ScenesabstractThe pedestrian-following robot aims to robustly maintain a standard distance from a specific pedestrian in an unknown complex environment, which is challenging when facing the limited field of view (FOV) and the computing resource constraints. However, the pedestrian-following robot commonly utilizes the long range dependence of targets to construct tracking models via the transformer tracking method, which is often insufficient to achieve real-time following. To address this important issue, we propose a new fast tracking network for the following robot, including a lightweight target detector, a target state prediction decoder, and an adaptive visual servo controller. First, in the lightweight target detector, a dense feature extraction process is designed by stacking depthwise group-separable convolutions. A context encoder is also developed, and is coupled with the proposed person re-identification (Re-ID) branch to detect multiple targets with high precision and speed. Second, in the target state prediction decoder, we introduce a flexible multiattention mechanism to obtain target Re-ID features from the previous frame for predicting the target's position in the current frame. Third, in the adaptive visual servo controller, we design a six-parameter dynamic model for the proportional integral derivative (PID) controller to stably follow the pedestrian under the limited FOV. Extensive experimental results demonstrate that the proposed method is efficient and accurate. Moreover, it also exhibits strong robustness and high real-time performance in unknown complex environments. Qin Wan 0001, Zhi Li 0087, Yaonan Wang 0001, Ruifeng Lv, Huaying Cheng, Di Wu 0046 |
IEEE Trans. Ind. Informatics | 3 |
| 2024 | A Physical-Constrained Decomposition Method of Infrared Thermography: Pseudo Restored Heat Flux Approach Based on Ensemble Bayesian Variance Tensor FractionabstractIn this study, we propose a new post processing algorithm, using a stable low-rank decomposed pseudo restored heat flux based on the ensemble variational Bayes tensor factorization (EVBTF-RPHF) algorithm for performing periodic square wave thermographic nondestructive testing (thermographic NDT). Previous studies have shown that both RPHF and EVBTF can separately improve the detectability of thermography by enhancing some defect features. However, both methods are limited by their particularly constraints: RPHF are heavily degraded by noises and missing data due to the assumptions under which the physical models are derived while efficiency of EVBT reduces when the lateral heat diffusion weights out. By embedding RPHF into the stable low-rank decomposition EVBTF, the proposed algorithm allows to improve the detectability of defects in thermographic NDT using a periodic heat flux with low-rank spatial distribution. The study verifies the capacity of the proposed method by theoretical analysis. Then, experiments were conducted on a carbon fiber composite panel with foreign inserts buried up to 5 mm deep. The sampled data are processed by the proposed method. The results are compared with existing methods such as phase-locked RPHF and EVBTF. The experimental results demonstrated that defects with normalized diameter-to-depth ratios as small as 0.9, barely detected with other available techniques, can reliably be detected by EVBTF-RPHF. The signal to noise ratio and the contrast are used as figure of merit to quantitatively compare the capacity of the proposed method with existing methods. However, the computation efficiency of the proposed algorithms needs further improvement. Hongjin Wang, Yuejun Hou, Yunze He, Can Wen, Benjamin Giron-Palomares, Yuxia Duan, Bin Gao 0003, Vladimir P. Vavilov, Yaonan Wang 0001 |
IEEE Trans. Ind. Informatics | 9 |
| 2024 | Data-Based Guaranteed Trajectory Estimation for Unmanned Surface VehiclesabstractThis article concerns the guaranteed trajectory estimation problem for unmanned surface vehicles (USVs) via set-membership estimation technique. Taking both rigid-body and hydrostatics kinetics into consideration, the nonlinear dynamic model of USV system is derived, where the parameters of system are all unknown. Considering external disturbance and nonlinearities, an offline data-based set-membership estimation algorithm of unknown system parameters is proposed to obtain the set representation of system parameters, which contain the actual parameters of USV. Then, based on the obtained parameter sets, an online guaranteed trajectory estimation algorithm of USV is constructed to provide guaranteed sets enclosing actual trajectory points of USV, which consists of a time-update step and measurement-update step. To tackle the nonlinear transformation of zonotopes, both interval arithmetic and Taylor model are utilized to provide rigorous bounds for nonlinear function. Finally, simulation results on a USV dynamic are provided to demonstrate the effectiveness of the proposed data-based guaranteed trajectory estimation method for USVs. Xudong Wang 0008, Hui Zhang 0023, Yuan Wang 0040, Yaonan Wang 0001 |
IEEE Trans. Ind. Informatics | 4 |
| 2024 | A Systematic Point Cloud Edge Detection Framework for Automatic Aircraft Skin MillingabstractThe edge detection technique is an essential step for aircraft skin milling in aviation manufacturing. Most of the current detection methods focus on traditionally defined edge extraction tasks but disregard the crucial systematic requirement of edge milling. In this article, we proposed a novel edge detection framework for automatic edge milling of aircraft skins. First, an edge probability detector is proposed by the spatial tangent continuity to provide the essential reference. Second, we propose a hierarchical branch searching method to hierarchically strip the desired milling edges from the raw point cloud, which consists of the following three graded progressive steps: branch backbone generation, branch extension, and branch pruning. We demonstrate the performance of the proposed method on both synthetic models and aircraft skin workpieces. The proposed method outperforms the other baselines and shows accurate edges for the edge milling task. Yaonan Wang 0001, He Xie, Mingtao Feng, Haotian Wu 0002, Chao Ding 0006, Ajmal Mian |
IEEE Trans. Ind. Informatics | 2 |
| 2024 | PoseDiffusion: A Coarse-to-Fine Framework for Unseen Object 6-DoF Pose EstimationabstractAccurately estimating the six-degrees of freedom (DoF) pose of unseen objects is crucial for successful robotic manipulation in industrial automation. Some existing methods for this task rely on prior knowledge of individual objects, i.e., the model must be trained on the exact object instance or object category. Others perform unseen object pose estimation but are limited in their feature learning and pose refinement ability. To address these problems, we propose an unseen object pose estimation method that follows a coarse-to-fine framework and leverages the powerful learning ability of diffusion models. We introduce a diffusion model for generating object poses, and conduct a comparison between the generated poses and the original pose to determine the optimal one. We design a novel pose estimation module to provide coarse poses for the PoseDiffusion. This module comprises two feature extraction modules that extract global and masked features. In addition, we propose a strategy to estimate the pose by comparing the similarity between rendered and query poses. The renderings of an unseen object from various viewpoints are generated from its computer-aided design (CAD) model. Our method requires a CAD model of the unseen object only during inference, a scenario well suited to industrial applications. Experimental evaluation on benchmark datasets demonstrates that the proposed framework outperforms existing approaches, achieving state-of-the-art performance in six-DoF object pose estimation. Qing Zhu 0003, Yaonan Wang 0001, Mingtao Feng, Chengzhong Wu, Xuebing Liu, Jianan Huang 0002, Ajmal Mian |
IEEE Trans. Ind. Informatics | 3 |
| 2024 | Flex-DLD: Deep Low-Rank Decomposition Model With Flexible Priors for Hyperspectral Image Denoising and RestorationabstractHyperspectral images (HSIs) are composed of hundreds of contiguous waveband images, offering a wealth of spatial and spectral information. However, the practical use of HSIs is often hindered by the presence of complicated noise caused by various factors such as non-uniform sensor response and dark current. Traditional methods for denoising HSIs rely on constrained optimization approaches, where selecting appropriate prior knowledge is critical for achieving satisfactory results. Nevertheless, these traditional algorithms are limited by hand-crafted priors, leaving room for improvement in their denoising performance. Recently, the supervised deep learning technique has emerged as a promising approach for HSI denoising. However, their requirement for paired training data and poor generalization ability on untrained noise distributions pose challenges in practical applications. In this paper, we design a novel algorithm by the synergism of optimization-based methods and deep learning techniques. Specifically, we introduce a plug-and-play Deep Low-rank Decomposition (DLD) model into the optimization framework. Furthermore, we propose an effective mechanism to incorporate traditional prior knowledge into the DLD model. Finally, we provide a detailed analysis of the optimization process and convergence of the proposed method. Empirical evaluations on various tasks, including hyperspectral image denoising and spectral compressive imaging, demonstrate the superiority of our approach over state-of-the-art methods. Yurong Chen 0003, Hui Zhang 0023, Yaonan Wang 0001, Yimin Yang 0001, Q. M. Jonathan Wu |
IEEE Trans. Image Process. | 3 |
| 2024 | MG-GCT: A Motion-Guided Graph Convolutional Transformer for Traffic Gesture RecognitionabstractFor autonomous driving systems, it is crucial to recognize the actions and gestures of traffic conductors and cyclists on the road to ensure safety. However, traffic gesture recognition is more challenging than action recognition in general scenarios due to the differences in action posture and sample composition between traffic gesture datasets and general action datasets. Therefore, general action recognition methods cannot identify traffic gestures well. To overcome these problems, we propose a novel motion-guided graph convolutional transformer (MG-GCT) for traffic gesture recognition. Firstly, we proposed a two-stream network to fully utilize joint data and motion data for action recognition. Secondly, we designed and implemented a motion-guided module between two streams, which leverages the powerful spatial representation ability of the motion data to guide the learning of the joint data stream in the spatial dimension. Thirdly, we implemented a temporal transformer network to process the temporal features of the skeleton. Finally, we conducted extensive experiments on two public datasets and one dataset presented by us to demonstrate the effectiveness of our network in traffic gesture recognition, which has a significant advantage over the state-of-the-art methods. Qing Zhu 0003, Yaonan Wang 0001, Yang Mo |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2024 | TransKD: Transformer Knowledge Distillation for Efficient Semantic SegmentationabstractSemantic segmentation benchmarks in the realm of autonomous driving are dominated by large pre-trained transformers, yet their widespread adoption is impeded by substantial computational costs and prolonged training durations. To lift this constraint, we look at efficient semantic segmentation from a perspective of comprehensive knowledge distillation and aim to bridge the gap between multi-source knowledge extractions and transformer-specific patch embeddings. We put forward the Transformer-based Knowledge Distillation (TransKD) framework which learns compact student transformers by distilling both feature maps and patch embeddings of large teacher transformers, bypassing the long pre-training process and reducing the FLOPs by >85.0%. Specifically, we propose two fundamental modules to realize feature map distillation and patch embedding distillation, respectively: 1) Cross Selective Fusion (CSF) enables knowledge transfer between cross-stage features via channel attention and feature map distillation within hierarchical transformers; 2) Patch Embedding Alignment (PEA) performs dimensional transformation within the patchifying process to facilitate the patch embedding distillation. Furthermore, we introduce two optimization modules to enhance the patch embedding distillation from different perspectives: 1) Global-Local Context Mixer (GL-Mixer) extracts both global and local information of a representative embedding; 2) Embedding Assistant (EA) acts as an embedding method to seamlessly bridge teacher and student models with the teacher’s number of channels. Experiments on Cityscapes, ACDC, NYUv2, and Pascal VOC2012 datasets show that TransKD outperforms state-of-the-art distillation frameworks and rivals the time-consuming pre-training method. The source code is publicly available athttps://github.com/RuipingL/TransKD. Ruiping Liu 0001, Kailun Yang 0001, Alina Roitberg, Jiaming Zhang 0001, Kunyu Peng, Huayao Liu, Yaonan Wang 0001, Rainer Stiefelhagen |
IEEE Trans. Intell. Transp. Syst. | 7 |
| 2024 | Imagery Overlap Block Compressive Sensing With Convex OptimizationabstractTo improve reconstruction performance in imagery compressive sensing, the present paper changes solving a block image compressive sensing reconstruction into a convex optimization problem. First, a Total-Variation norm minimization constraints model that includes both L1 and L2 norm functions is established. The split Bregman iterative method solves the model with convex optimization. Then, a robust adaptive image block compressive sensing algorithm is studied based on an analysis of the image features. The image is divided into blocks, and an overlap image block compressive reconstruction method is proposed. Finally, to solve the block effect caused by block compressive sensing reconstruction, a novel image overlap block compressive sensing reconstruction based on the Poisson function is suggested to avoid the block effect in the reconstruction process. The experimental results show that compared with other traditional compressive sensing reconstruction algorithms, the proposed method can generate a better image reconstruction result. According to the PSNR evaluation, when the sampling rate is 0.3, the proposed method is improved by more than 20.98% compared to the conventional techniques, and according to the SSIM evaluation, it has improved by more than 11.92% from the traditional methods. We can also find that the proposed method has better construction effect for traffic sign image recognition compared with ordinary natural image reconstruction. When the sampling rate is only 0.1, the PSNR value reaches 44.28dB, and the SSIM reconstruction accuracy reaches 98.14%. After reconstructing different types and characteristic images, it is supported that the proposed algorithm has good robustness and anti-noise performance. Lin Zhang 0041, Yudong Zhang 0001, Yaonan Wang 0001 |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2024 | LSKANet: Long Strip Kernel Attention Network for Robotic Surgical Scene SegmentationabstractSurgical scene segmentation is a critical task in Robotic-assisted surgery. However, the complexity of the surgical scene, which mainly includes local feature similarity (e.g., between different anatomical tissues), intraoperative complex artifacts, and indistinguishable boundaries, poses significant challenges to accurate segmentation. To tackle these problems, we propose the Long Strip Kernel Attention network (LSKANet), including two well-designed modules named Dual-block Large Kernel Attention module (DLKA) and Multiscale Affinity Feature Fusion module (MAFF), which can implement precise segmentation of surgical images. Specifically, by introducing strip convolutions with different topologies (cascaded and parallel) in two blocks and a large kernel design, DLKA can make full use of region- and strip-like surgical features and extract both visual and structural information to reduce the false segmentation caused by local feature similarity. In MAFF, affinity matrices calculated from multiscale feature maps are applied as feature fusion weights, which helps to address the interference of artifacts by suppressing the activations of irrelevant regions. Besides, the hybrid loss with Boundary Guided Head (BGH) is proposed to help the network segment indistinguishable boundaries effectively. We evaluate the proposed LSKANet on three datasets with different surgical scenes. The experimental results show that our method achieves new state-of-the-art results on all three datasets with improvements of 2.6%, 1.4%, and 3.4% mIoU, respectively. Furthermore, our method is compatible with different backbones and can significantly increase their segmentation accuracy. Code is available at https://github.com/YubinHan73/LSKANet. Min Liu 0008, Yubin Han, Jiazheng Wang 0001, Can Wang 0011, Yaonan Wang 0001, Erik Meijering |
IEEE Trans. Medical Imaging | 5 |
| 2024 | Brain Image Segmentation for Ultrascale Neuron Reconstruction via an Adaptive Dual-Task Learning NetworkabstractAccurate morphological reconstruction of neurons in whole brain images is critical for brain science research. However, due to the wide range of whole brain imaging, uneven staining, and optical system fluctuations, there are significant differences in image properties between different regions of the ultrascale brain image, such as dramatically varying voxel intensities and inhomogeneous distribution of background noise, posing an enormous challenge to neuron reconstruction from whole brain images. In this paper, we propose an adaptive dual-task learning network (ADTL-Net) to quickly and accurately extract neuronal structures from ultrascale brain images. Specifically, this framework includes an External Features Classifier (EFC) and a Parameter Adaptive Segmentation Decoder (PASD), which share the same Multi-Scale Feature Encoder (MSFE). MSFE introduces an attention module named Channel Space Fusion Module (CSFM) to extract structure and intensity distribution features of neurons at different scales for addressing the problem of anisotropy in 3D space. Then, EFC is designed to classify these feature maps based on external features, such as foreground intensity distributions and image smoothness, and select specific PASD parameters to decode them of different classes to obtain accurate segmentation results. PASD contains multiple sets of parameters trained by different representative complex signal-to-noise distribution image blocks to handle various images more robustly. Experimental results prove that compared with other advanced segmentation methods for neuron reconstruction, the proposed method achieves state-of-the-art results in the task of neuron reconstruction from ultrascale brain images, with an improvement of about 49% in speed and 12% in F1 score. Min Liu 0008, Shuhan Wu, Zhuangdian Lin, Yaonan Wang 0001, Erik Meijering |
IEEE Trans. Medical Imaging | 5 |
| 2024 | R2D2-GAN: Robust Dual Discriminator Generative Adversarial Network for Microscopy Hyperspectral Image Super-ResolutionabstractHigh-resolution microscopy hyperspectral (HS) images can provide highly detailed spatial and spectral information, enabling the identification and analysis of biological tissues at a microscale level. Recently, significant efforts have been devoted to enhancing the resolution of HS images by leveraging high spatial resolution multispectral (MS) images. However, the inherent hardware constraints lead to a significant distribution gap between HS and MS images, posing challenges for image super-resolution within biomedical domains. This discrepancy may arise from various factors, including variations in camera imaging principles (e.g., snapshot and push-broom imaging), shooting positions, and the presence of noise interference. To address these challenges, we introduced a unique unsupervised super-resolution framework named R2D2-GAN. This framework utilizes a generative adversarial network (GAN) to efficiently merge the two data modalities and improve the resolution of microscopy HS images. Traditionally, supervised approaches have relied on intuitive and sensitive loss functions, such as mean squared error (MSE). Our method, trained in a real-world unsupervised setting, benefits from exploiting consistent information across the two modalities. It employs a game-theoretic strategy and dynamic adversarial loss, rather than relying solely on fixed training strategies for reconstruction loss. Furthermore, we have augmented our proposed model with a central consistency regularization (CCR) module, aiming to further enhance the robustness of the R2D2-GAN. Our experimental results show that the proposed method is accurate and robust for super-resolution images. We specifically tested our proposed method on both a real and a synthetic dataset, obtaining promising results in comparison to other state-of-the-art methods. Hui Zhang 0023, Jiang-Huai Tian, Yingjian Su, Yurong Chen 0003, Yaonan Wang 0001 |
IEEE Trans. Medical Imaging | 6 |
| 2024 | Occlusion-Aware Feature Recover Model for Occluded Person Re-IdentificationabstractOccluded person re-identification (Re-ID) is a challenging task, as various object-to-person (OTP) and person-to-person (PTP) occlusion scenarios cause diverse occlusion interference and target person feature loss problems in person matching. Most existing methods, which utilize auxiliary models to evaluate the unoccluded person parts for occlusion feature elimination, are inefficient and cannot handle the PTP occlusion scenarios and person feature loss problems. To solve these issues, we propose a novel Occlusion-Aware Feature Recover (OAFR) model. OAFR simulates diverse occlusions to facilitate the model perceiving OTP, PTP occlusions and recovers occluded query features with unoccluded retrieved gallery features. Concretely, the Prior Knowledge-based Occlusion Simulation method is firstly introduced to synthesize OTP, PTP occlusions and corresponding occlusion labels, empowering model target person perception and occlusion-aware capability through self-supervised learning. Afterward, the feature recovery module reconstructs occluded query features with corresponding unoccluded local features of the top-$K$retrieved images by the visibility weighted average scheme, thus recovering the occluded query features to maintain more comprehensive features for better retrieval. Extensive experiments demonstrate that the proposed OAFR achieves superior performance to the state-of-the-art for both holistic and occluded Re-ID. Especially for Occluded-DukeMTMC dataset, OAFR outperforms the state-of-the-art by 6.0% for Rank-1 accuracy and 2.2% for mAP score. The source codes are available athttps://github.com/yuanbianGit/OAFR. Yuan Bian 0002, Min Liu 0008, Yaonan Wang 0001 |
IEEE Trans. Multim. | 5 |
| 2024 | Relation-Preserving Feature Embedding for Unsupervised Person Re-IdentificationabstractSome unsupervised approaches have been proposed recently for the person re-identification (ReID) problem since annotations of samples across cameras are time-consuming. However, most of these methods focus on the appearance content of the sample itself, and thus seldom take the structure relations among samples into account when learning the feature representation, which would provide a valuable guide for learning the representations of the samples. Thus hard samples may not be well solved due to the limited or even misleading information of the sample itself. To address this issue, in this article, we propose a Relation-Preserving Feature Embedding (RPE) model that leverages structure relations among samples to boost the performance of the unsupervised person ReID methods without requiring any sample annotations. RPE aims at integrating the sample content and the neighborhood structure relations among samples into the learning of feature embeddings by combining the advantages of the autoencoder and graph autoencoder. Specifically, a relation and content information fusion (RCIF) module is proposed to dynamically merge the information from both perspectives of content and relation levels for feature embedding learning. Also, due to the lack of the identity labels of samples, we adopt an adaptive optimization strategy to update the affinity relations among samples instead of the reconstruction of the whole affinity matrix for optimizing the RPE model, which is more suitable for the unsupervised ReID task. Rigorous experiments on three widely-used large-scale benchmarks for person ReID demonstrate the superiority of the proposed method over current state-of-the-art unsupervised methods. Min Liu 0008, Fei Wang 0124, Jianhua Dai 0003, Anan Liu, Yaonan Wang 0001 |
IEEE Trans. Multim. | 6 |
| 2024 | Multi-Layer Decoupling Attention Network for Weakly Supervised Object LocalizationabstractWeakly supervised object localization (WSOL) aims to localize the entire and well-defined objects only via image-level labels for reducing the need of labor-intensive annotation and mitigating annotation errors. However, many WSOL methods via class activation maps (CAMs) often suffer from incomplete activation and inaccurate boundaries for object localization. In this article, we propose a novel multi-layer decoupling attention localization (MDAL) network to address these issues. We first present a simple yet effective multi-layer comparison decoupling mechanism including a maximum decoupling function and a minimum decoupling function to sufficiently activate and fuse multi-layer features. Then, we introduce the multi-layer maximum decoupling function into the attention modules, and develop a channel attention activation decoupling (CAAD) module and a spatial attention activation decoupling (SAAD) module, which can mine much more useful information for more possible regions' activation. Furthermore, the multi-layer minimum decoupling function is introduced to efficiently fuse and refine multi-layer features, which can suppress the over-activation and background noise. Finally, we develop a joint loss function to train the MDAL network. Experimental results on CUB-200-2011 and ILSVRC2012 demonstrate that our proposed network can provide accurate and complete object localization. Aoran Zhang 0002, Zhigang Ling, Yaonan Wang 0001 |
IEEE Trans. Multim. | 3 |
| 2024 | Modified Noise-Immune Fuzzy Neural Network for Solving the Quadratic Programming With Equality Constraint ProblemabstractQuadratic programming with equality constraint (QPEC) problems have extensive applicability in many industries as a versatile nonlinear programming modeling tool. However, noise interference is inevitable when solving QPEC problems in complex environments, so research on noise interference suppression or elimination methods is of great interest. This article proposes a modified noise-immune fuzzy neural network (MNIFNN) model and use it to solve QPEC problems. Compared with the traditional gradient recurrent neural network (TGRNN) and traditional zeroing recurrent neural network (TZRNN) models, the MNIFNN model has the advantage of inherent noise tolerance ability and stronger robustness, which is achieved by combining proportional, integral, and differential elements. Furthermore, the design parameters of the MNIFNN model adopt two disparate fuzzy parameters generated by two fuzzy logic systems (FLSs) related to the residual and residual integral term, which can improve the adaptability of the MNIFNN model. Numerical simulations demonstrate the effectiveness of the MNIFNN model in noise tolerance. Jianhua Dai 0003, Liu Luo, Lin Xiao 0002, Lei Jia 0001, Penglin Cao, Jichun Li 0002, Natalio Krasnogor, Yaonan Wang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 8 |
| 2024 | Event-Triggered Adaptive Neural Impedance Control of Robotic SystemsabstractThis article presents an event-triggered adaptive neural impedance control (ETANIC) scheme for robotic systems, where the combination of impedance control (IC) and event-triggered mechanism can significantly reduce the computational burden and the communication cost under the premise of ensuring the stability and tracking performances of the robotic systems. The IC is used to achieve the compliant behavior of the robotic systems in response to the environment. The uncertainties of the robotic systems are estimated by the radial basis function neural network (RBFNN), and the update laws for RBFNN are derived from the designed Lyapunov function. The stability of the whole closed-loop control system is analyzed by the Lyapunov theory, and the event-triggered conditions are designed to avoid the Zeno behavior. The numerical simulation and experimental tests demonstrate that the proposed ETANIC scheme can achieve better efficiency for controlling the robotic systems to perform the interaction tasks with the environment in comparison to the adaptive neural IC (ANIC). Shuai Ding 0007, Jinzhu Peng, Hui Zhang 0023, Yaonan Wang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2024 | SwinPA-Net: Swin Transformer-Based Multiscale Feature Pyramid Aggregation Network for Medical Image SegmentationabstractThe precise segmentation of medical images is one of the key challenges in pathology research and clinical practice. However, many medical image segmentation tasks have problems such as large differences between different types of lesions and similar shapes as well as colors between lesions and surrounding tissues, which seriously affects the improvement of segmentation accuracy. In this article, a novel method called Swin Pyramid Aggregation network (SwinPA-Net) is proposed by combining two designed modules with Swin Transformer to learn more powerful and robust features. The two modules, named dense multiplicative connection (DMC) module and local pyramid attention (LPA) module, are proposed to aggregate the multiscale context information of medical images. The DMC module cascades the multiscale semantic feature information through dense multiplicative feature fusion, which minimizes the interference of shallow background noise to improve the feature expression and solves the problem of excessive variation in lesion size and type. Moreover, the LPA module guides the network to focus on the region of interest by merging the global attention and the local attention, which helps to solve similar problems. The proposed network is evaluated on two public benchmark datasets for polyp segmentation task and skin lesion segmentation task as well as a clinical private dataset for laparoscopic image segmentation task. Compared with existing state-of-the-art (SOTA) methods, the SwinPA-Net achieves the most advanced performance and can outperform the second-best method on the mean Dice score by 1.68%, 0.8%, and 1.2% on the three tasks, respectively. Jiazheng Wang 0001, Min Liu 0008, Yaonan Wang 0001, Erik Meijering |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2024 | A Signed Subgraph Encoding Approach via Linear Optimization for Link Sign PredictionabstractIn this article, we consider the problem of inferring the sign of a link based on known sign data in signed networks. Regarding this link sign prediction problem, signed directed graph neural networks (SDGNNs) provides the best prediction performance currently to the best of our knowledge. In this article, we propose a different link sign prediction architecture called subgraph encoding via linear optimization (SELO), which obtains overall leading prediction performances compared to the state-of-the-art algorithm SDGNN. The proposed model utilizes a subgraph encoding approach to learn edge embeddings for signed directed networks. In particular, a signed subgraph encoding approach is introduced to embed each subgraph into a likelihood matrix instead of the adjacency matrix through a linear optimization (LO) method. Comprehensive experiments are conducted on five real-world signed networks with area under curve (AUC), F1, micro-F1, and macro-F1 as the evaluation metrics. The experiment results show that the proposed SELO model outperforms existing baseline feature-based methods and embedding-based methods on all the five real-world networks and in all the four evaluation metrics. Zhihong Fang, Shaolin Tan, Yaonan Wang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2024 | Computation-Efficient Fault Detection Framework for Partially Known Nonlinear Distributed Parameter SystemsabstractFault detection for distributed parameter systems (DPSs) generally requires the complete model information to be known so far. However, for numerous industrial applications, it is common that accurate first-principles physical models are extremely difficult to obtain. Hence, the applicability of traditional model-based methods is being restricted. To pave the way, an adaptive neural network (AdNN) is constructed to simultaneously estimate the state variable and the unknown nonlinearity for a class of partially known nonlinear DPSs. Moreover, considering that full-state measurement is unrealistic in applications, the proposed adaptive neural observer is based on a reduced-order model, which also increases the computation efficiency. Then, the residual generation and evaluation are conducted using the output estimation error of the proposed adaptive neural observer. Bearing the effects of the neglected fast dynamics in mind, a data-driven threshold generation scheme is proposed. Extensive experimental results are presented and analyzed to validate the effectiveness of the proposed method. Yun Feng 0001, Yaonan Wang 0001, Yang Mo, Yiming Jiang 0001, Zhijie Liu 0001, Wei He 0001, Han-Xiong Li |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2024 | Modal-Regression-Based Broad Learning System for Robust Regression and ClassificationabstractA novel neural network, namely, broad learning system (BLS), has shown impressive performance on various regression and classification tasks. Nevertheless, most BLS models may suffer serious performance degradation for contaminated data, since they are derived under the least-squares criterion which is sensitive to noise and outliers. To enhance the model robustness, in this article we proposed a modal-regression-based BLS (MRBLS) to tackle the regression and classification tasks of data corrupted by noise and outliers. Specifically, modal regression is adopted to train the output weights instead of the minimum mean square error (MMSE) criterion. Moreover, the$\ell_{2,1}$-norm-induced constraint is used to encourage row sparsity of the connection weight matrix and achieve feature selection. To effectively and efficiently train the network, the half-quadratic theory is used to optimize MRBLS. The validity and robustness of the proposed method are verified on various regression and classification datasets. The experimental results demonstrate that the proposed MRBLS achieves better performance than the existing state-of-the-art BLS methods in terms of both accuracy and robustness. Licheng Liu, Tingyun Liu, C. L. Philip Chen, Yaonan Wang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2024 | A Timestamp-Based Inertial Best-Response Dynamics for Distributed Nash Equilibrium Seeking in Weakly Acyclic GamesabstractIn this article, we consider the problem of distributed game-theoretic learning in games with finite action sets. A timestamp-based inertial best-response dynamics is proposed for Nash equilibrium seeking by players over a communication network. We prove that if all players adhere to the dynamics, then the states of players will almost surely reach consensus and the joint action profile of players will be absorbed into a Nash equilibrium of the game. This convergence result is proven under the condition of weakly acyclic games and strongly connected networks. Furthermore, to encounter more general circumstances, such as games with graphical action sets, state-based games, and switching communication networks, several variants of the proposed dynamics and its convergent results are also developed. To demonstrate the validity and applicability, we apply the proposed timestamp-based learning dynamics to design distributed algorithms for solving some typical finite games, including the coordination games and congestion games. Shaolin Tan, Zhihong Fang, Yaonan Wang 0001, Jinhu Lü 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2024 | A Dynamic-Varying Parameter Enhanced ZNN Model for Solving Time-Varying Complex-Valued Tensor Inversion With Its Application to Image EncryptionabstractTime-varying complex-valued tensor inverse (TVCTI) is a public problem worthy of being studied, while numerical solutions for the TVCTI are not effective enough. This work aims to find the accurate solution to the TVCTI using zeroing neural network (ZNN), which is an effective tool in terms of solving time-varying problems and is improved in this article to solve the TVCTI problem for the first time. Based on the design idea of ZNN, an error-adaptive dynamic parameter and a new enhanced segmented signum exponential activation function (ESS-EAF) are first designed and applied to the ZNN. Then a dynamic-varying parameter-enhanced ZNN (DVPEZNN) model is proposed to solve the TVCTI problem. The convergence and robustness of the DVPEZNN model are theoretically analyzed and discussed. In order to highlight better convergence and robustness of the DVPEZNN model, it is compared with four varying-parameter ZNN models in the illustrative example. The results show that the DVPEZNN model has better convergence and robustness than the other four ZNN models in different situations. In addition, the state solution sequence generated by the DVPEZNN model in the process of solving the TVCTI cooperates with the chaotic system and deoxyribonucleic acid (DNA) coding rules to obtain the chaotic-ZNN-DNA (CZD) image encryption algorithm, which can encrypt and decrypt images with good performance. Lin Xiao 0002, Penglin Cao, Yongjun He 0001, Wensheng Tang, Jichun Li 0002, Yaonan Wang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 7 |
| 2024 | Anchor Association Learning for Unsupervised Video Person Re-IdentificationabstractVideo-based person re-identification (re-id) has attracted a significant attention in recent years due to the increasing demand of video surveillance. However, existing methods are usually based on the supervised learning, which requires vast labeled identities across cameras and is not suitable for real scenes. Although some unsupervised approaches have been proposed for video re-id, their performance is far from satisfactory. In this article, we propose an unsupervised anchor association learning (UAAL) framework to address the video-based person re-id task, in which the feature representation of each sampled tracklet is regarded as an anchor. Specifically, we first propose an intracamera anchor association learning (IAAL) term that learns the discriminative anchor by utilizing the affiliation relations between an image and the anchors in each camera. Then, the exponential moving average (EMA) strategy is employed to update the anchor and the updated anchors are stored into an anchor memory module. On top of that, a cross-camera anchor association learning (CAAL) term is introduced to mine potential positive anchor pairs across cameras by presenting a cyclic ranking anchor alignment and threshold filtering method. Extensive experiments conducted on two public datasets show the superiority of the proposed method; for example, our method achieves 73.2% for rank-1 accuracy and 60.1% for mean average precision (mAP) score, respectively, on MARS, similarly 89.7% and 87.0% on DukeMTMC-VideoReID. Shujun Zeng, Min Liu 0008, Qing Liu 0035, Yaonan Wang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2024 | Reinforcement Learning-Based Adaptive Optimal Control for Nonlinear Systems With Asymmetric HysteresisabstractThis article investigates the adaptive optimal tracking problem for a class of nonlinear affine systems with asymmetric Prandtl-Ishlinskii (PI) hysteresis nonlinearities based on actor-critic (A-C) learning mechanisms. Considering the huge obstacles arising from the uncertainty of hysteresis nonlinearity in actuators, we develop a scheme for the conflict between the construction of Hamilton functions and hysteresis nonlinearity. The actuator hysteresis forces the input into a hysteresis delay, thus preventing the Hamilton function from getting the current moment's input instantly and thus making optimization impossible. In the first step, an inverse model is constructed to compensate for the hysteresis model with a shift factor. In the second step, we compensate for the control input by designing a feedback controller and incorporating the estimation and approximation errors into the Hamilton error. Optimal control, the other part of the actual control input, is obtained by taking partial derivatives of the Hamiltonian function after the nonlinearities have been circumvented. At the end, a simulation is given to validate the developed solution. Licheng Zheng, Zhi Liu 0001, Yaonan Wang 0001, C. L. Philip Chen, Yun Zhang 0001, Zongze Wu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2024 | PDE Model-Based On-Line Cell-Level Thermal Fault Localization Framework for BatteriesabstractUnknown distributed incipient thermal fault detection and localization are vital to the safe operation of batteries while they have not been given sufficient attention in existing works compared to the studies on estimation of State of Charge (SoC) as well as State of Health (SoH). In order to fill this gap, a backstepping-based fault localization filter (FLF) is presented. Generally, full-state temperature measurement is required to achieve fault localization, which is impossible in applications. However, with the help of interpolation-based approximation, the required number of sensors decreases from infinity to only a few, which guarantees the usability of FLF. A comprehensive methodology framework, including the FLF design, residual evaluation in a distributed manner, and threshold computation, is introduced to guarantee reliable and robust performance in a real-time pattern. Theoretic analysis as well as experiment validations are presented to for validation. Yun Feng 0001, Yaonan Wang 0001, Bing-Chuan Wang, Hui Zhang 0023, Zhengguang Wu, Huaicheng Yan 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 3 |
| 2024 | Distributed Neuro-Adaptive Prescribed Performance Consensus Control for Nontriangular Structural Multiagent SystemsabstractFor a class of nontriangular structural multiagent systems, this article presents a neuro-adaptive prescribed performance consensus control scheme. By using mean value theorem to isolate the virtual variables and neural networks to approximate the ideal controller, the system model is reconstructed, based on which the virtual controllers are able to be derived. The algebraic-loop problem is circumvented utilizing the properties of basis function. With the proposed performance functions, an error transformation is presented, based on which the controller scheme is developed. It is ensured that the follower agents synchronize at a predefined speed, and synchronization error converges to a specified range within a given time. Two Simulations demonstrate the effectiveness of the presented control method. Zhuangbi Lin, Junhe Liu, Yaonan Wang 0001, C. L. Philip Chen, Yun Zhang 0001, Zhi Liu 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 3 |
| 2024 | Cyclic-Small-Gain Approach to Adaptive Control for Multiagent Systems With Unknown Interconnected DynamicsabstractDeveloping a distributed output feedback consensus tracking control scheme for nonlinear multiagent systems (MASs) with interconnected dynamics and unmeasurable states holds significant practical importance. Current approaches to this challenge often rely on fuzzy logic systems (FLSs) or neural networks (NNs) for direct compensation of interconnected terms. However, these methods frequently result in incomplete compensation, posing obstacles to achieving true distributed control. In this study, the cyclic-small-gain condition is utilized to address the challenge of unknown interconnected dynamics. This approach effectively decouples the physical coupling among MASs, enabling the realization of distributed control. Further, direct compensation using the cyclic-small-gain condition introduces the potential singularity problem, which we effectively overcome by employing the inequality technique. Based on the presented control scheme, the existing control results for the MASs with interconnected dynamics are extended from stabilization control to trajectory tracking. As proved, the synchronization and observer errors eventually converge to a tunable zero region, enabling each agent to effectively track the leader. Additionally, the proposed control scheme employs a direct adaptive law design approach with a significantly reduced number of adaptive laws. Simulation results validate the theoretical scheme. Meijian Tan, Zhi Liu 0001, Yaonan Wang 0001, C. L. Philip Chen, Yun Zhang 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 3 |
| 2024 | Comprehensive Study on Zeroing Neural Network With High-Order Evolutionary Formula, Nonlinear Functions, and Variable Parameter for Time-Changing Matrix Cholesky DecompositionabstractIn this article, a low-order zeroing neural network (LZNN), a high-order ZNN (HZNN), and a variable-parameter ZNN (VZNN) are designed and applied to the time-changing Cholesky decomposition of any positive-definite matrix, where the LZNN and HZNN models are generated based on the traditional and high-order evolutionary formulas, respectively. In addition, a new activation function (N-Acf) is applied to the LZNN, HZNN, and VZNN models to improve the convergence and robustness. Importantly, the LZNN and HZNN models activated by the N-Acf have faster predefined-time convergence velocity when solving the time-changing Cholesky decomposition problem of any positive-definite matrix, which is demonstrated via theoretical analysis and numerical experiments. Finally, in light of empirical and theoretical evidence, it can be established that the solution model of the VZNN model is able to undergo convergence to the theoretical solution of Cholesky decomposition despite the presence of interposing noise. Lin Xiao 0002, Sida Xiao, Yongjun He 0001, Jianhua Dai 0003, Yaonan Wang 0001, Yiwei Li 0006 |
IEEE Trans. Syst. Man Cybern. Syst. | 5 |
| 2024 | Resilient Formation Control With Koopman Operator for Networked NMRs Under Denial-of-Service AttacksabstractThis article presents a resilient formation control framework for networked nonholonomic mobile robots (NMRs) that enables long-time recovery abilities subject to denial-of-service (DoS) attacks by taking advantage of the Koopman operator. Due to the intermittent interruption of communication under DoS, the transmitted signals among the networked NMRs are incomplete. In the lifted space, the infinite-dimensional Koopman operator is employed to capture a linear characteristic of the missed signals from the available signals. Specifically, a data-driven cost function is developed to approximate the infinite-dimensional Koopman operator, allowing long-term recovery capabilities for the missed signals, where the useful historical data is identified by an event-triggered mechanism (ETM). Then, the least-squares method is implemented to calculate a finite-dimensional approximation of the Koopman operator. Once DoS attacks are active, the missed signals are recovered forward from the latest received signals through the approximation Koopman operator. Furthermore, according to the recovered and transmitted signals, the resilient formation controller with a variable gain takes into account the convergence rate and the steady state formation error. The Lyapunov theorem is introduced to prove that the formation error quickly converges to the minor compact set. A distributed DoS attack example is conducted to validate the efficiency and superiority in numerical simulation, and the proposed method is implemented on the real networked NMRs. Weiwei Zhan, Zhiqiang Miao, Hui Zhang 0023, Zhengguang Wu, Wei He 0001, Yaonan Wang 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 7 |
| 2023 | Sketch and Text Guided Diffusion Model for Colored Point Cloud GenerationabstractDiffusion probabilistic models have achieved remarkable success in text guided image generation. However, generating 3D shapes is still challenging due to the lack of sufficient data containing 3D models along with their descriptions. Moreover, text based descriptions of 3D shapes are inherently ambiguous and lack details. In this paper, we propose a sketch and text guided probabilistic diffusion model for colored point cloud generation that conditions the denoising process jointly with a hand drawn sketch of the object and its textual description. We incrementally diffuse the point coordinates and color values in a joint diffusion process to reach a Gaussian distribution. Colored point cloud generation thus amounts to learning the reverse diffusion process, conditioned by the sketch and text, to iteratively recover the desired shape and color. Specifically, to learn effective sketch-text embedding, our model adaptively aggregates the joint embedding of text prompt and the sketch based on a capsule attention network. Our model uses staged diffusion to generate the shape and then assign colors to different parts conditioned on the appearance prompt while preserving precise shapes from the first stage. This gives our model the flexibility to extend to multiple tasks, such as appearance re-editing and part segmentation. Experimental results demonstrate that our model outperforms recent state-of-the-art in point cloud generation. Yaonan Wang 0001, Mingtao Feng, He Xie, Ajmal Mian |
ICCV | 2 |
| 2023 | VDBblox: Accurate and Efficient Distance Fields for Path Planning and Mesh ReconstructionabstractHighly accurate and efficient map in unknown and complex environments is essential for robotics navigation. Traditionally, mobile robot platforms are often computationally constrained when using multiple sensors to process large amounts of input data. In previous works, some of them have been deployed to embedded platforms in real-time. However, how to balance accuracy and efficiency while reducing the computational resources and the memory footprint is still the bottleneck. Motivated by these challenges, we proposed a mapping framework called VDBblox to incrementally build Euclidean Signed Distance Fields (ESDFs) map from Truncated Signed Distance Fields (TSDFs) mapping. We use a novel weight function to update the non-projective TSDFs, thus improving the quality of the mesh reconstruction with higher accuracy than up-to-date methods. Meanwhile, the generated ESDFs map is maintained by the least recently used (LRU) cache to dynamically handle the obstacle changes with less runtime than state-of-the-art. We show VDBblox performance in terms of accuracy and efficiency by benchmark comparison on RGB-D and LiDAR public datasets. Moreover, we demonstrate that VDBblox can be integrated into a completed quadrotor system as a sub-module. Then we validate it through online obstacle avoidance and high-quality mesh reconstruction in real-world experiments. Finally, we release our method as open-source code to the community11Code - https://github.com/yinloonga/vdbblox. Yinlong Bai, Zhiqiang Miao, Xiangke Wang, Yong Liu 0007, Hesheng Wang 0001, Yaonan Wang 0001 |
IROS | 6 |
| 2023 | Deep Stereo Matching with Superpixel Based Feature and Cost
Kai Zeng 0010, Hui Zhang 0023, Wei Wang 0025, Yaonan Wang 0001, Jianxu Mao |
PRCV (2) | 4 |
| 2023 | Improved YOLOv7 Based on Transformer for Object Detection in UAV-Captured ImagesabstractAs the drone captures image targets at different flying altitudes, their scales may vary significantly, which can pose challenges for the object detection model to accurately detect them. Additionally, tiny objects in the image contain minimal information, making them difficult to distinguish from the background. To overcome these two challenges, we proposed a network architecture that aims to improve the accuracy of tiny object detection in drone images. Specially, we designed a tiny object detector(TOD) that can effectively extract features of tiny objects and distinguish between tiny object features and image background. Furthermore, this TOD module contains a Convolutional Visual Attention Network (CVAN) to better focus on the regions of tiny objects. Experimental results demonstrate that the proposed method achieves [email protected] accuracy of 53.9% on the VisDrone2021-test-dev dataset and improves by 2.8 % compared to YOLOv7. Yuefan Luo, Qing Zhu 0003, Zhen Zhou 0003, Lin Chen 0034, Tianjian Jiang, Yijiang Li, Danwei Wang, Yaonan Wang 0001 |
SMC | 9 |
| 2023 | Separable-programming based probabilistic-iteration and restriction-resolving correlation filter for robust real-time visual tracking
Baiheng Cao, Xuedong Wu, Jianxu Mao, Yaonan Wang 0001 |
Eng. Appl. Artif. Intell. | 4 |
| 2023 | Group consensus for competitive multiagent systems with input saturation and intermittent communication using current and outdated statesabstractThis paper investigates group consensus control for first-order multi-agent systems (MASs) with input saturation and intermittent communication. A novel class of distributed group consensus protocols is designed by combining current and outdated states of agents, which differs from the standard one using only current states. Sufficient conditions for group consensus are derived for MASs without two conservative assumptions required by previous literature. The proposed protocol can accelerate the convergence speed of MASs by choosing appropriate outdated states. Several examples are provided to illustrate the effectiveness of our results. Jiacheng Su, Yaonan Wang 0001 |
Neurocomputing | 3 |
| 2023 | Aggregated decentralized down-sampling-based ResNet for smart healthcare systems
Zhiwen Jiang, Ziji Ma, Yaonan Wang 0001, Xun Shao, Keping Yu, Alireza Jolfaei |
Neural Comput. Appl. | 3 |
| 2023 | Prior Image Guided Snapshot Compressive Spectral ImagingabstractSpectral images with rich spatial and spectral information have wide usage, however, traditional spectral imaging techniques undeniably take a long time to capture scenes. We consider the computational imaging problem of the snapshot spectral spectrometer, i.e., the Coded Aperture Snapshot Spectral Imaging (CASSI) system. For the sake of a fast and generalized reconstruction algorithm, we propose a prior image guidance-based snapshot compressive imaging method. Typically, the prior image denotes the RGB measurement captured by the additional uncoded panchromatic camera of the dual-camera CASSI system. We argue that the RGB image as a prior image can provide valuable semantic information. More importantly, we design the Prior Image Semantic Similarity (PIDS) regularization term to enhance the reconstructed spectral image fidelity. In particular, the PIDS is formulated as the difference between the total variation of the prior image and the recovered spectral image. Then, we solve the PIDS regularized reconstruction problem by the Alternating Direction Method of Multipliers (ADMM) optimization algorithm. Comprehensive experiments on various datasets demonstrate the superior performance of our method. Yurong Chen 0003, Yaonan Wang 0001, Hui Zhang 0023 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2023 | Adaptive Prescribed Performance Control of Unmanned Aerial Manipulator With DisturbancesabstractThis article presents the problem of autonomous control of an unmanned aerial manipulator (UAM) developed for operation with unknown disturbances, wherein the disturbances from the coupling effect between the UAM and the external environment need to be considered. Regarding the coupling force as a disturbance to the entire UAM system, an adaptive prescribed performance control (APPC) scheme utilizing the knowledge of prescribed performance is proposed to guarantee the transient and steady-state performance responses. Also, an adaptive law is designed to estimate the upper boundary parameters of the UAM system uncertainties and disturbances, wherein the restrictive constant boundary assumptions and the prior information of the upper bound are not required in the controller design. Furthermore, to enable safe manipulation in a realistic situation, an end-effector trajectory generation method is presented satisfying the joint angle limitation. For the validation of the proposed method, the simulation results of numerical simulation comparisons are shown. Moreover, experimental scenarios including stable flight and simulated co-work with humans in complex environments are designed to verify the proposed method.Note to Practitioners—This article is motivated by the problem of aerial manipulation under unknown disturbances, which may be caused by the wide movement of the manipulator and the sudden loading or unloading of an object. Existing approaches for aerial manipulation often require the assumption of a constant or slowly varying external disturbance. However, a priori bounded disturbance might impose a priori bound on the system state before obtaining closed-loop stability. In this article, the proposed controller with an adaptive law is designed to estimate the upper boundary parameters of the overall disturbances and ensure the predefined performance, so that the prior information of the upper bound of disturbances is not required. The performance of the proposed control strategy is demonstrated via numerical simulation comparisons and experiments, including stale flight and simulated co-work with humans in a complex environment. Jiacheng Liang, Yangning Wu, Zhiqiang Miao, Hui Zhang 0023, Yaonan Wang 0001 |
IEEE Trans Autom. Sci. Eng. | 6 |
| 2023 | Robust Image-Based Landing Control of a Quadrotor on an Unpredictable Moving Vehicle Using Circle FeaturesabstractThis paper addresses the landing problem of a quadrotor on an unpredictable moving vehicle, using a robust image-based visual servoing (IBVS) method. The circle-based image moments are defined to construct image dynamics, and the passivity-like property of the circle features is preserved by reprojecting to the virtual image plane. The landing control system is decoupled into translation and rotation modules due to the rotation invariance of the proposed circle features. First, by exploiting the error transformation in the image space, a robust IBVS controller can overcome the lack of both the desired depth information of target features and the velocity feedback of the target. Next, an adaptive geometric attitude controller is developed directly using rotation matrices to avoid the singularities of Euler-angles and the ambiguity of quaternions. One benefit of the proposed scheme is that it can potentially improve the camera visibility, guarantee the transient and steady-state behaviors in image space, and be efficiently implemented on the low-cost quadrotor. Finally, The stability analysis is presented using Lyapunov stability theory on cascaded systems, and the effectiveness of the proposed control strategy is demonstrated through simulations and experiments. Note to Practitioners—The motivation of this paper is to investigate a practical control strategy for the image-based landing control of underactuated quadrotors on an unpredictable moving vehicle. In most of the existing image-based landing control schemes for underactuated quadrotors, having the prior predictive model of the moving landing vehicle to provide a feed-forward compensation during the landing maneuver. However, due to the fact that the landing environment and vehicle are primarily stochastic, resulting in no predictive models are valid in practice. Therefore, this paper suggests a robust image-based landing control strategy without the model or state of the moving landing vehicle. In particular, a novel virtual circle feature, possessing the characteristic of rotation invariance, is designed for the landing of underactuated quadrotors, which decouples the landing system and simplifies the control design. Moreover, the image feature errors are directly retained within prescribed performance funnels in the image space. As a result, the transient and steady-state landing behaviors can be implicitly guaranteed in Cartesian space. The stability and convergence of the system are analyzed mathematically and the experiment using quadrotors provides promising results. In ongoing research, we are addressing the issues of collision avoidances and unknown disturbances to provide a more realistic setup for the autonomous deployment and recovery of underactuated quadrotors in GPS-denied environments. Jie Lin 0009, Yaonan Wang 0001, Zhiqiang Miao, Hesheng Wang 0001, Rafael Fierro |
IEEE Trans Autom. Sci. Eng. | 2 |
| 2023 | Robust Guaranteed Synchronization for Chaotic Systems With Incremental Quadratic ConstraintsabstractThis paper concerns the robust guaranteed synchronization of chaotic systems subject to unknown but amplitude-bounded parameter perturbation, process disturbance, and measurement noise. The incremental quadratic constraints are adopted to provide a unified description of chaotic systems, based on which a circle criterion-based observer with$\ell _{\infty }$performance is design to attenuate the impacts of unknown uncertainties on synchronization error. Based on the designed observer, a zonotope-based guaranteed synchronization algorithm is proposed via mean value extension and zonotope inclusion techniques, where all the uncertainties from parameter perturbation, process disturbance, measurement noise, and chaotic behavior are taken into consideration. Finally, simulation results on the hyper-chaotic Chen system are presented to demonstrate the effectiveness of the proposed method. Xudong Wang 0008, Guoqi Wang, Zhe Li 0050, Yaonan Wang 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 4 |
| 2023 | Adaptive Refining-Aggregation-Separation Framework for Unsupervised Domain Adaptation Semantic SegmentationabstractUnsupervised domain adaptation has attracted widespread attention as a promising method to solve the labeling difficulties of semantic segmentation tasks. It trains a segmentation network for unlabeled real target images using easily available labeled virtual source images. To improve performance, clustering is used to obtain domain-invariant feature representations. However, most clustering-based methods indiscriminately cluster all features mapped by category from both domains, causing the centroid shift and affecting the generation of discriminative features. We propose a novel clustering-based method that uses an adaptive refining-aggregation-separation framework, which learns the discriminative features by designing different adaptive schemes for different domains and features. The clustering does not require any tunable thresholds. To estimate more accurate domain-invariant centroids, we design different ways to guide the adaptive refinement of different domain features. A critic is proposed to directly evaluate the confidence of target features to solve the absence of target labels. We introduce a domain-balanced aggregation loss and two adaptive separation losses for distance and similarity respectively, which can discriminate clustering features by combining the refinement strategy to improve segmentation performance. Experimental results on GTA$5\rightarrow $Cityscapes and SYNTHIA$\rightarrow $Cityscapes benchmarks show that our method outperforms existing state-of-the-art methods. Yihong Cao, Hui Zhang 0023, Xiao Lu 0002, Yurong Chen 0003, Yaonan Wang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2023 | SNIS: A Signal Noise Separation-Based Network for Post-Processed Image Forgery DetectionabstractImage forgery detection has aroused widespread research interest in both academia and industry because of its potential security threats. Existing forgery detection methods achieve excellent tampered regions localization performance when forged images have not undergone post-processing, which can be detected by observing changes in the statistical features of images. However, forged images may be carefully post-processed to conceal forgery boundaries in a particular scenario. It becomes tough challenging to these methods. In this paper, we perform an analogous analysis between image forgery detection and blind signal separation, and formulate the post-processed image forgery detection problem into a signal noise separation problem. We also propose a signal noise separation-based (SNIS) network to solve the problem of detecting post-processed image forgery. Specifically, we first adopt the signal noise separation module to separate tampered region from the complex background region with post-processing noise, which weakens or even eliminates the negative impact of post-processing on forgery detection. Then, the multi-scale feature learning module uses a parallel atrous convolution architecture to learn high-level global features from multiple perspectives. Besides, a feature fusion module is utilized to enhance the discriminability of tampered regions and real regions by strengthening the boundary information. Finally, the prediction module is designed to predict the tampered region and classify the type of tampering operation. Extensive experiments show that the proposed SNIS is not only effective for forgery detection on forged images without post-processing, but also promising in robustness against multiple post-processing attacks. Furthermore, SNIS is robust in detecting forged images from unknown sources. Xin Liao 0001, Wei Wang 0025, Zhenxing Qian, Zheng Qin 0001, Yaonan Wang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2023 | Adaptive Sliding-Mode Disturbance Observer-Based Finite-Time Control for Unmanned Aerial Manipulator With Prescribed PerformanceabstractIn this article, an adaptive sliding-mode disturbance observer (ASMDO)-based finite-time control scheme with prescribed performance is proposed for an unmanned aerial manipulator (UAM) under uncertainties and external disturbances. First, to take into account the dynamic characteristics of the UAM, a dynamic model of the UAM with state-dependent uncertainties and external disturbances is introduced. Then, note that a priori bounded uncertainty may impose a priori constraint on the system state before obtaining closed-loop stability. To remove this assumption, an ASMDO with a nested adaptive structure is introduced to effectively estimate and compensate the external disturbances and state-dependent uncertainties in finite time without the information of the upper bound of the uncertainties and disturbances and their derivatives. Furthermore, based on the proposed ASMDO, the finite-time control scheme with the prescribed performance is presented to ensure finite-time convergence and implement the specified transient and steady-state performance. The Lyapunov tools are utilized to analyze the stability of the proposed controller. Finally, the correctness and performance of the proposed controller are illustrated through numerical simulation comparisons and outdoor experimental comparisons. Jiacheng Liang, Yangning Wu, Zhiqiang Miao, Hui Zhang 0023, Yaonan Wang 0001 |
IEEE Trans. Cybern. | 6 |
| 2023 | D-BIN: A Generalized Disentangling Batch Instance Normalization for Domain AdaptationabstractPattern recognition is significantly challenging in real-world scenarios by the variability of visual statistics. Therefore, most existing algorithms relying on the independent identically distributed assumption of training and test data suffer from the poor generalization capability of inference on unseen testing datasets. Although numerous studies, including domain discriminator or domain-invariant feature learning, are proposed to alleviate this problem, the data-driven property and lack of interpretation of their principle throw researchers and developers off. Consequently, this dilemma incurs us to rethink the essence of networks' generalization. An observation that visual patterns cannot be discriminative after style transfer inspires us to take careful consideration of the importance of style features and content features. Does the style information related to the domain bias? How to effectively disentangle content and style features across domains? In this article, we first investigate the effect of feature normalization on domain adaptation. Based on it, we propose a novel normalization module to adaptively leverage the propagated information through each channel and batch of features called disentangling batch instance normalization (D-BIN). In this module, we explicitly explore domain-specific and domaininvariant feature disentanglement. We maneuver contrastive learning to encourage images with the same semantics from different domains to have similar content representations while having dissimilar style representations. Furthermore, we construct both self-form and dual-form regularizers for preserving the mutual information (MI) between feature representations of the normalization layer in order to compensate for the loss of discriminative information and effectively match the distributions across domains. D-BIN and the constrained term can be simply plugged into state-of-the-art (SOTA) networks to improve their performance. In the end, experiments, including domain adaptation and generalization, conducted on different datasets have proven their effectiveness. Yurong Chen 0003, Hui Zhang 0023, Yaonan Wang 0001, Weixing Peng, Wangdong Zhang, Q. M. Jonathan Wu, Yimin Yang 0001 |
IEEE Trans. Cybern. | 3 |
| 2023 | Self-Paced Broad Learning SystemabstractBroad learning system (BLS), an efficient neural network with a flat structure, has received a lot of attention due to its advantages in training speed and network extensibility. However, the conventional BLS adopts the least square loss, which treats each sample equally and thus is sensitivity to noise and outliers. To address this concern, in this article we propose a self-paced BLS (SPBLS) model by incorporating the novel self-paced learning (SPL) strategy into the network for noisy data regression. With the assistance of the SPL criterion, the model output is used as feedback to learn appropriate priority weight to readjust the importance of each sample. Such a reweighting strategy can help SPBLS to distinguish samples from "easy" to "difficult" in model training, equipping the model robust to noise and outliers while maintaining the characteristics of the original system. Moreover, two incremental learning algorithms associated to SPBLS have also been developed, with which the system can be updated quickly and flexibly without retraining the entire model when new training samples are added or the network needs to be expanded. Experiments conducted on various datasets demonstrate that the proposed SPBLS can achieve satisfying performance for noisy data regression. Licheng Liu, Luyang Cai, Ting Xie 0003, Yaonan Wang 0001 |
IEEE Trans. Cybern. | 4 |
| 2023 | An Augmented Game Approach for Design and Analysis of Distributed Learning Dynamics in Multiagent GamesabstractIn this article, an augmented game approach is proposed for the formulation and analysis of distributed learning dynamics in multiagent games. Through the design of the augmented game, the coupling structure of utility functions among all the players can be reformulated into an arbitrary undirected connected network while the Nash equilibria are preserved. In this case, any full-information game learning dynamics can be recast into a distributed form, and its convergence can be determined from the structure of the augmented game. We apply the proposed approach to generate both deterministic and stochastic distributed gradient play and obtain several negative convergent results about the distributed gradient play: 1) a Nash equilibrium is convergent under the classic gradient play, yet its corresponding augmented Nash equilibrium may be not convergent under the distributed gradient play and, on the other side, 2) a Nash equilibrium is not convergent under the classic gradient play, yet its corresponding augmented Nash equilibrium may be convergent under the distributed gradient play. In particular, we show that the variational stability structure (including monotonicity as a special case) of a game is not guaranteed to be preserved in its augmented game. These results provide a systematic methodology about how to formulate and then analyze the feasibility of distributed game learning dynamics. Shaolin Tan, Zhihong Fang, Yaonan Wang 0001, Jinhu Lü 0001 |
IEEE Trans. Cybern. | 3 |
| 2023 | Distributed Bayesian Inference Over Sensor NetworksabstractIn this article, two novel distributed variational Bayesian (VB) algorithms for a general class of conjugate-exponential models are proposed over synchronous and asynchronous sensor networks. First, we design a penalty-based distributed VB (PB-DVB) algorithm for synchronous networks, where a penalty function based on the Kullback-Leibler (KL) divergence is introduced to penalize the difference of posterior distributions between nodes. Then, a token-passing-based distributed VB (TPB-DVB) algorithm is developed for asynchronous networks by borrowing the token-passing approach and the stochastic variational inference. Finally, applications of the proposed algorithm on the Gaussian mixture model (GMM) are exhibited. Simulation results show that the PB-DVB algorithm has good performance in the aspects of estimation/inference ability, robustness against initialization, and convergence speed, and the TPB-DVB algorithm is superior to existing token-passing-based distributed clustering algorithms. Baijia Ye, Jiahu Qin, Weiming Fu, Yingda Zhu, Yaonan Wang 0001, Yu Kang 0001 |
IEEE Trans. Cybern. | 5 |
| 2023 | Intensive Noise-Tolerant Zeroing Neural Network Based on a Novel Fuzzy Control ApproachabstractTo overcome the disadvantages of the current zeroing neural network (ZNN) in noise tolerance, this article first proposes an intensive noise-tolerant ZNN (INT-ZNN) by introducing a novel fuzzy control approach (FCA). This FCA is designed dexterously according to the variation of two errors related to the INT-ZNN. Thus, the most feature of the INT-ZNN is that the added fuzzy control can inherently restrain the various noises. Compared with the previous noise-tolerant ZNN derived by the integral design formula, the INT-ZNN with a much simpler structure can tolerate the noise in finite/fixed time. That is, the INT-ZNN activated by nonlinear functions possesses finite/fixed-time convergence while suppressing the noise, which is guaranteed by the presented theorems. Besides, it also theoretically proves that the INT-ZNN has global stability under the interference of noise. In the simulative experiment, the INT-ZNN is used to solve the time-varying Sylvester matrix equation problem and the experimental results verify the excellent noise-tolerance of the INT-ZNN. Meanwhile, the INT-ZNN is successfully applied to image processing. Lei Jia 0001, Lin Xiao 0002, Jianhua Dai 0003, Yaonan Wang 0001 |
IEEE Trans. Fuzzy Syst. | 4 |
| 2023 | Adaptive Fuzzy Prescribed Performance Output-Feedback Cooperative Control for Uncertain Nonlinear Multiagent SystemsabstractThis article considers the prescribed performance output-feedback cooperative control of multiagent systems with nonparametric uncertainties. Scale and performance functions are proposed to construct a barrier function, based on which the adaptive prescribed performance control scheme is developed. Therein, the fuzzy observers are employed to estimate the unavailable states. The proposed scheme has the following characteristics: first, the tracking errors are always limited to a specified boundary. Second, the tracking error converges to a specified accuracy within a given time, and the settling time and tracking accuracy are determined only by the user and no longer depend on the initial conditions. Third, the distributed fuzzy controller design uses only the output information from agent and its neighbor. The effectiveness and practicality of the proposed method are demonstrated by two simulation examples. Zhuangbi Lin, Zhi Liu 0001, Chun-Yi Su, Yaonan Wang 0001, C. L. Philip Chen, Yun Zhang 0001 |
IEEE Trans. Fuzzy Syst. | 4 |
| 2023 | Adaptive Fuzzy Inverse Optimal Control of Nonlinear Switched SystemsabstractExisting approaches to optimal control of uncertain switched systems require heavy computations for approximating and learning the optimal solution of the Hamilton–Jacobi–Bellman equation. To overcome this problem, a fuzzy adaptive inverse optimal control strategy is first developed for switched systems, which minimizes not only the cost function but also circumvents solving the Hamilton–Jacobi–Bellman equation. Specifically, an alternative practical inverse approach is developed by two Lyapunov functions method, based on which the inverse optimality regarding a meaningful cost function is achieved. In addition, a new condition of admissible edge-dependent average dwell time is developed by applying the extended multiple Lyapunov functions method. Guided by this weaker condition, the stability of the considered switched system is proved. Finally, simulations are carried out to verify the developed method. Danping Zeng, Zhi Liu 0001, Yaonan Wang 0001, C. L. Philip Chen, Yun Zhang 0001, Zongze Wu 0001 |
IEEE Trans. Fuzzy Syst. | 3 |
| 2023 | A Fast Online Planning Under Partial Observability Using Information Entropy RewardsabstractMotion planning in an unknown environment is a common challenge because of the existing uncertainties. Representatively, the partially observable Markov decision process (POMDP) is a general mathematical framework for planning in uncertain environments. Recent POMDP solvers generally adopt the sparse reward scheme to solve the planning under uncertainty problem. Subsequently, the robot's exploration may be hindered without immediate rewards, resulting in excessively long planning time. In this article, a POMDP method, information entropy determinized sparse partially observation tree (IE-DESPOT), is proposed to explore a high-quality solution and efficient planning in unknown environments. First, a novel sample method integrating state distribution and Gaussian distribution is proposed to optimize the quality of the sampled states. Then, an information entropy based on sampled states is established for real-time reward calculation, resulting in the improvement of robot exploration efficiency. Moreover, the near-optimality and convergence of the proposed algorithm are analyzed. As a result, compared with general-purpose POMDP solvers, the proposed algorithm exhibits fast convergence to a near-optimal policy in many examples of interest. Furthermore, the IE-DESPOT's performance is verified in real mobile robot experiments. Jiangjiang Liu 0005, Limin Lan, Hui Zhang 0023, Zhiqiang Miao, Yaonan Wang 0001 |
IEEE Trans. Ind. Informatics | 6 |
| 2023 | Transformer-Based Imitative Reinforcement Learning for Multirobot Path PlanningabstractMultirobot path planning leads multiple robots from start positions to designated goal positions by generating efficient and collision-free paths. Multirobot systems realize coordination solutions and decentralized path planning, which is essential for large-scale systems. The state-of-the-art decentralized methods utilize imitation learning and reinforcement learning methods to teach fully decentralized policies, dramatically improving their performance. However, these methods cannot enable robots to perform tasks efficiently in relatively dense environments without communication between robots. We introduce the transformer structure into policy neural networks for the first time, dramatically enhancing the ability of policy neural networks to extract features that facilitate collaboration between robots. It mainly focuses on improving the performance of policies in relatively dense multirobot environments under conditions where robots do not communicate with each other. Furthermore, a novel imitation reinforcement learning framework is proposed by combining contrastive learning and double deep Q-network to solve the problem of difficulty training policy neural networks after introducing the transformer structure. We present results in the simulation environment and compare the resulting policy against advanced multirobot path-planning methods in terms of success rate. Simulation results show that our policy achieves state-of-the-art performance when there is no communication between robots. Finally, we experimented with a real-world case using a total of three robots in our robotic laboratory. Lin Chen 0034, Yaonan Wang 0001, Zhiqiang Miao, Yang Mo, Mingtao Feng, Zhen Zhou 0003, Hesheng Wang 0001 |
IEEE Trans. Ind. Informatics | 2 |
| 2023 | Image-Based Visual Servoing of Unmanned Aerial Manipulators for Tracking and Grasping a Moving TargetabstractIn this article, an image-based visual servoing (IBVS) control strategy is proposed for the unmanned aerial manipulator (UAM) system to track and grasp a moving target. Specifically, a robust-adaptive velocity observer is designed to estimate the relative velocity between the tracked target and the UAM platform. Based on the velocity observer, an IBVS controller using onboard camera of the UAM platform is proposed for moving target tracking without velocity measurement. Then, the barrier Lyapunov function is introduced into the UAM platform IBVS controller to ensure the safety of target tracking. Besides, another virtual camera is constructed on manipulator end-effector to compensate for the tracking error of the UAM platform. As a benefit, the eye-to-hand onboard camera ensures the global view of the UAM, and the eye-in-hand virtual camera of the manipulator ensures the accuracy of the grasping task. Finally, the stability of the proposed IBVS control strategy is analyzed through Lyapunov theory. The comparative simulations are provided to illustrate the target tracking performance of the proposed method. The experimental results demonstrate that the proposed method can be applied to the UAM with a low-cost sensor suite to realize the tasks of tracking and grasping a moving target. Yangning Wu, Zhiqiang Miao, Hang Zhong, Hui Zhang 0023, Yaonan Wang 0001 |
IEEE Trans. Ind. Informatics | 7 |
| 2023 | SET: Sampling-Enhanced Exploration Tree for Mobile Robot in Restricted EnvironmentsabstractMobile robots generally work in harsh and restricted environments, which poses challenges for mobile robots to find a feasible path efficiently. This article presents a planning method, namely, sampling-enhanced exploration tree (SET), to improve computational efficiency in restricted environments while guaranteeing high-quality performance. The core of SET is sampling-enhanced exploration, which consists of critical areas identification, guiding-exploration, and rectifying-exploration. In the critical areas identification phase, the restricted areas are identified based on the distribution of the hybrid samples. Next, the critical samples in restricted areas are selected as the origins of the sampling-enhanced exploration. In the guiding-exploration phase, the sampling-enhanced exploration starts from the origins and marches quickly with the guidance of the leader-samples to capture the spatial feature and connectivity of the restricted areas. The spatial information provides essential guidance for efficient biased sampling. In the rectifying-exploration phase, the directions of sampling-enhanced exploration are rectified to transit the problematic areas and supplement samples. Theoretical analysis is provided to shed light on the properties of SET. Moreover, the generality and effectiveness of SET are verified through a series of mobile robot simulations and real-world experiments. Zhiqiang Miao, Hui Zhang 0023, Yaonan Wang 0001 |
IEEE Trans. Ind. Informatics | 6 |
| 2023 | SSAPN: Spectral-Spatial Anomaly Perception Network for Unsupervised Vaccine DetectionabstractVaccines are the most significant and effective way to prevent disease and safeguard human health. However, it is easy to produce or mix foreign matters during the manufacturing process. Moreover, foreign matters are extremely faint that it is difficult to obtain images and detect them accurately. To tackle imaging challenges, in this article, we built a hyperspectral imaging system to construct a first-of-its-kind HSI dataset with pixel-level annotation for vaccine anomaly detection, where the vaccine comes from the actual pharmaceutical company. To address the problem of low detection accuracy, we propose a spectral–spatial anomaly perception module joint with an unsupervised autoencoder network (SSAPN), in which nonlinear features learned from the encoder are divided into nonoverlapping patches and mapped to efficiently encode spectral and spatial feature information. The spectral–spatial multilayer perceptrons (MLP) module consists of continuous and alternating spectral MLP with spatial MLP, which achieves spectral with spatial perception in the global receptive field, captures long-range dependencies, and extracts the most discriminative spectral–spatial features. Experimental results show that our SSAPN model outperforms other state-of-the-art anomaly detection methods in terms of both detection and generalization performance. This work will help speed up the production process in the vaccine pharmaceutical industry and ensure vaccine quality. Ating Yin, Yaonan Wang 0001, Yurong Chen 0003, Kai Zeng 0010, Hui Zhang 0023, Jianxu Mao |
IEEE Trans. Ind. Informatics | 2 |
| 2023 | OTP-NMS: Toward Optimal Threshold Prediction of NMS for Crowded Pedestrian DetectionabstractPedestrian detection is still a challenging task for computer vision, especially in crowded scenes where the overlaps between pedestrians tend to be large. The non-maximum suppression (NMS) plays an important role in removing the redundant false positive detection proposals while retaining the true positive detection proposals. However, the highly overlapped results may be suppressed if the threshold of NMS is lower. Meanwhile, a higher threshold of NMS will introduce a larger number of false positive results. To solve this problem, we propose an optimal threshold prediction (OTP) based NMS method that predicts a suitable threshold of NMS for each human instance. First, a visibility estimation module is designed to obtain the visibility ratio. Then, we propose a threshold prediction subnet to determine the optimal threshold of NMS automatically according to the visibility ratio and classification score. Finally, we re-formulate the objective function of the subnet and utilize the reward-guided gradient estimation algorithm to update the subnet. Comprehensive experiments on CrowdHuman and CityPersons show the superior performance of the proposed method in pedestrian detection, especially in crowded scenes. Min Liu 0008, Baopu Li, Yaonan Wang 0001, Wanli Ouyang |
IEEE Trans. Image Process. | 4 |
| 2023 | FedCrack: Federated Transfer Learning With Unsupervised Representation for Crack DetectionabstractEmpowered by labeled datasets, supervised pre-training based transfer learning (SPTL) has made significant advances for image classification applications. However, due to privacy-preserving protocol and unaccessible annotation, it emerges as a novel problem in federated learning scenarios whether unsupervised pre-training based transfer learning (UPTL) is available for semantic segmentation. In this work, we define federated transfer learning (FedTL) in the absence of source domain label, and track research progress on pavement crack benchmark. The main challenges of FedTL include: i) a privacy-protecting distributed training framework that extends UPTL to the constraints of federated settings, and ii) a self-learning semantic segmentation approach that develops self-supervised learning paradigm to simultaneously learn category and shape representations. Motivated by that, we propose a FedCrack model to absorb feature disentanglement and prototype clustering into vision Transformer, which obtains the pre-trained encoder on source domain without accessing annotation. Thereafter, a fine-tuning stage is presented to learn decoder with scaling attention on target domain for fine-grained crack segmentation. The effectiveness of proposed FedCrack can be demonstrated with superior performance of 82.14% on mIoU and 9.85 FPS on speed in extensive experiments. To the best of our knowledge, it is the first work in FedTL to gain weights of unsupervised pre-training representations on source domain locally, gradients of which are then aggregated to a federated central model that also fine-tunes the transferable parameters by target domain. Xiating Jin, Jiajun Bu, Hui Zhang 0023, Yaonan Wang 0001 |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2023 | Adaptive Dynamic Path Planning Method for Autonomous Vehicle Under Various Road Friction and SpeedsabstractPath planning is a crucial technology for autonomous vehicle (AV). However, it is difficult to adapt to dynamic driving environment, and AV may lose lateral dynamic stability due to high speed and various friction. This paper presents an adaptive dynamic path planning method (ADPPM) for AV to address the challenges. The ADPPM is comprised of three components: 1) A dynamic state-fused steering decision method based on a hierarchical fuzzy inference system is designed to calculate the steering position of AV in each iteration step, and the method fuses the multi-state of the dynamic environment; 2) dynamic path optimization method is designed to reduce the mean curvature of the path based on the particle swarm optimization method, which improves the lateral dynamics stability of AV; 3) adaptive speed inference method is proposed to provide desired steering speed for AV according to various road friction and reduce AV’s steering burden. The ADPPM provides path planning in the dynamic environment, and it also improves the stability of AV under various road friction and speeds. Finally, the proposed method is verified by CarSim. Xiaofang Yuan, Zhixian Liu, Weihua Tan, Xizheng Zhang, Yaonan Wang 0001 |
IEEE Trans. Intell. Transp. Syst. | 6 |
| 2023 | Deep Confidence Propagation Stereo NetworkabstractStereo matching depth estimation based on rectified image pairs is of great importance to many computer vision tasks such as vehicle navigation and autonomous driving. Confidence measures are typically used to refine stereo matching results, which provides robustness and efficiency for disparity estimation. However, previous learning-based confidence methods for stereo matching usually use the middle results or composition as a post-processing step to refine the stereo matching results. This cannot be optimized end-to-end and the performance is limited by the quality of the tri-modal output. To handle this issue, in this paper, we pursue an end-to-end hierarchical architecture and propose a differentiable confidence propagation (DCP) model of a cost aggregation network for stereo matching. The DCP model is integrated into an end-to-end neural network hierarchical architecture to guide matching cost volume aggregation. More specifically, to better represent the similarity of left and right feature maps, we extract unary context feature maps with an effective attention mechanism for matching cost construction. Moreover, we aggregate the cost volume with the multiple stacked DCP cost aggregation (DCPCA) networks to generate a more reliable and finer cost volume. This network suppresses multi-level disparity maps. Each output disparity is supervised with different training weights to learn in a coarse-to-fine way. Our method outperforms previous methods on the Sceneflow dataset by achieving the$0.6735px$EPE error, achieving 1.53% D1-all metric of Non-occluded pixels regions and 0.72% Non-occluded pixels of$5px$metric on KITTI 2015 and 2012 dataset. Extensive experiments carried out on the KITTI Stereo benchmarks demonstrate that our DCPCA-Net can significantly minimize the trade-off between accuracy and efficiency for stereo matching. Kai Zeng 0010, Yaonan Wang 0001, Wei Wang 0025, Hui Zhang 0023, Jianxu Mao, Qing Zhu 0003 |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2023 | Elementary Subgraph Features for Link Prediction With Neural NetworksabstractThe enclosing subgraph of a target link has been proved to be effective for prediction of potential links. However, it is still unclear what topological features of the subgraph play the key role in determining the existence of links. To give a possible answer to this question, in this paper, we propose a neural network based learning method for link prediction with only 1-hop neighborhood information. In detail, we extract the one-hop neighborhood of a target link as the enclosing subgraph, then encode the subgraph into different types of topological features, and lastly feed these features to train a fully connected neural network for link prediction. The experimental results show that our proposed learning method with the 1-hop neighborhood features could outperform those heuristic-based methods and achieve nearly equal performance to the state-of-the-art learning-based method WLNM and SEAL. Furthermore, it is observed that these features can be concatenated with attribute vectors to greatly promote the link prediction performance in attributed graphs. This indicates that the topological pattern within an enclosing subgraph, which determines the existence of a possible link, can be aggregated by some elementary subgraph features. Zhihong Fang, Shaolin Tan, Yaonan Wang 0001, Jinhu Lü 0001 |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2023 | Branch Aggregation Attention Network for Robotic Surgical Instrument SegmentationabstractSurgical instrument segmentation is of great significance to robot-assisted surgery, but the noise caused by reflection, water mist, and motion blur during the surgery as well as the different forms of surgical instruments would greatly increase the difficulty of precise segmentation. A novel method called Branch Aggregation Attention network (BAANet) is proposed to address these challenges, which adopts a lightweight encoder and two designed modules, named Branch Balance Aggregation module (BBA) and Block Attention Fusion module (BAF), for efficient feature localization and denoising. By introducing the unique BBA module, features from multiple branches are balanced and optimized through a combination of addition and multiplication to complement strengths and effectively suppress noise. Furthermore, to fully integrate the contextual information and capture the region of interest, the BAF module is proposed in the decoder, which receives adjacent feature maps from the BBA module and localizes the surgical instruments from both global and local perspectives by utilizing a dual branch attention mechanism. According to the experimental results, the proposed method has the advantage of being lightweight while outperforming the second-best method by 4.03%, 1.53%, and 1.34% in mIoU scores on three challenging surgical instrument datasets, respectively, compared to the existing state-of-the-art methods. Code is available at https://github.com/SWT-1014/BAANet. Wenting Shen, Yaonan Wang 0001, Min Liu 0008, Jiazheng Wang 0001, Renjie Ding, Zhe Zhang 0022, Erik Meijering |
IEEE Trans. Medical Imaging | 2 |
| 2023 | 3D Soma Detection in Large-Scale Whole Brain Images via a Two-Stage Neural Networkabstract3D soma detection in whole brain images is a critical step for neuron reconstruction. However, existing soma detection methods are not suitable for whole mouse brain images with large amounts of data and complex structure. In this paper, we propose a two-stage deep neural network to achieve fast and accurate soma detection in large-scale and high-resolution whole mouse brain images (more than 1TB). For the first stage, a lightweight Multi-level Cross Classification Network (MCC-Net) is proposed to filter out images without somas and generate coarse candidate images by combining the advantages of the multi convolution layer's feature extraction ability. It can speed up the detection of somas and reduce the computational complexity. For the second stage, to further obtain the accurate locations of somas in the whole mouse brain images, the Scale Fusion Segmentation Network (SFS-Net) is developed to segment soma regions from candidate images. Specifically, the SFS-Net captures multi-scale context information and establishes a complementary relationship between encoder and decoder by combining the encoder-decoder structure and a 3D Scale-Aware Pyramid Fusion (SAPF) module for better segmentation performance. The experimental results on three whole mouse brain images verify that the proposed method can achieve excellent performance and provide the reconstruction of neurons with beneficial information. Additionally, we have established a public dataset named WBMSD, including 798 high-resolution and representative images ( 256 ×256 ×256 voxels) from three whole mouse brain images, dedicated to the research of soma detection, which will be released along with this paper. Xiaodan Wei, Qinghao Liu, Min Liu 0008, Yaonan Wang 0001, Erik Meijering |
IEEE Trans. Medical Imaging | 4 |
| 2023 | Refining Noisy Labels With Label Reliability Perception for Person Re-IdentificationabstractMost person re-identification (Re-ID) approaches rely excessively on a great quantity of annotated training data. However, due to sampling errors or annotated errors, the label noise is unavoidable, which usually causes a dramatic decrease in the performance of existing Re-ID methods. To address this problem, we propose the label reliability perception (LRP) for person Re-ID by refining noisy labels. Specifically, a feature-fusion block (FFB) is proposed to enhance the discrimen- ability of pedestrians' features by expanding the network's attention span due to the fused feature, which is generated by overlapping the coarse-grained feature obtained by global average pooling and fine-grained features obtained by evenly dividing the feature map in the height dimension and performing global max pooling. In addition, the label dual perception (LDP) is proposed to refine noisy labels instead of filtering samples by evaluating the reliability of each training sample's label. Specifically, we meticulously design five evaluation modes for each sample to perceive the reliability of the labels of thek-nearest neighbor images. Finally, we utilize the most reliable label to replace the noisy label and optimize the network. Extensive experiments prove the superiority of the proposed model over the competing methods; for instance, on Market1501, our method achieves 88.8% rank-1 accuracy and 70.5% mAP (4.7% and 4.3% improvements over the state-of-the-arts) under noise ratio 20%, and similarly on DukeMTMC-ReID, our method achieves 77.7% and 60.3%. Yongchun Chen, Min Liu 0008, Fei Wang 0124, Anan Liu, Yaonan Wang 0001 |
IEEE Trans. Multim. | 6 |
| 2023 | Cycle Consistency Based Pseudo Label and Fine Alignment for Unsupervised Domain AdaptationabstractUnsupervised Domain Adaptation (UDA) aims to transfer knowledge from a well-labeled source domain to an unlabeled target domain with a correlative distribution. Numerous existing approaches process this hard nut by directly matching the marginal distribution between two domains, which confront the obstacle of rough alignment and blurred decision boundary. Recent advances in UDA introduce target pseudo-label and subdomain adaptation to reduce misalignment and distribution discrepancy. Whereas, they frequently ignore that the production of target pseudo-label is so dependent on the source-trained classifier, which without reasonable restriction to discriminate generated pseudo-label is whether confident. Meanwhile, many methods in the subdomain alignment metric ignore exploring the potential distribution discrepancy between same-class samples of the intra-domain. To address these two issues simultaneously, this paper proposes a Cycle Consistency based Pseudo Label and Fine Alignment (CCPLFA) approach for UDA. In particular, firstly, a novel cycle-consistency based pseudo label module is designed, which is a simple yet effective way to alleviate the noise of pseudo labels and improve their semantic correctness. Secondly, we develop a Fine-Alignment distribution matching metric. Which can maximize the feature distribution density of intra-class cross-domains and not overlook the distribution structure of the global aspect. Comprehensive experiment results on four benchmarks demonstrate the capability of plug and play and the well generalization performance of our proposed method. Hui Zhang 0023, Junkun Tang, Yihong Cao, Yurong Chen 0003, Yaonan Wang 0001, Q. M. Jonathan Wu |
IEEE Trans. Multim. | 5 |
| 2023 | Design and Analysis of a Self-Adaptive Zeroing Neural Network for Solving Time-Varying Quadratic ProgrammingabstractIn order to solve the time-varying quadratic programming (TVQP) problem more effectively, a new self-adaptive zeroing neural network (ZNN) is designed and analyzed in this article by using the Takagi-Sugeno fuzzy logic system (TSFLS) and thus called the Takagi-Sugeno (T-S) fuzzy ZNN (TSFZNN). Specifically, a multiple-input-single-output TSFLS is designed to generate a self-adaptive convergence factor to construct the TSFZNN model. In order to obtain finite- or predefined-time convergence, four novel activation functions (AFs) [namely, power-bi-sign AF (PBSAF), tanh-bi-sign AF (TBSAF), exp-bi-sign AF (EBSAF), and sinh-bi-sign AF (SBSAF)] are developed and applied in the TSFZNN model for solving the TVQP problem. Both theoretical proofs and experimental simulations show that the TSFZNN model using PBSAF or TBSAF has the property of converging in a finite time, and the TSFZNN model using EBSAF or SBSAF has the property of converging in a predefined time, which have superior convergence performance compared to the traditional ZNN model. Jianhua Dai 0003, Lin Xiao 0002, Lei Jia 0001, Xinwang Liu 0002, Yaonan Wang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2023 | Modal Regression-Based Graph Representation for Noise Robust Face HallucinationabstractManifold learning-based face hallucination technologies have been widely developed during the past decades. However, the conventional learning methods always become ineffective in noise environment due to the least-square regression, which usually generates distorted representations for noisy inputs they employed for error modeling. To solve this problem, in this article, we propose a modal regression-based graph representation (MRGR) model for noisy face hallucination. In MRGR, the modal regression-based function is incorporated into graph learning framework to improve the resolution of noisy face images. Specifically, the modal regression-induced metric is used instead of the least-square metric to regularize the encoding errors, which admits the MRGR to robust against noise with uncertain distribution. Moreover, a graph representation is learned from feature space to exploit the inherent typological structure of patch manifold for data representation, resulting in more accurate reconstruction coefficients. Besides, for noisy color face hallucination, the MRGR is extended into quaternion (MRGR-Q) space, where the abundant correlations among different color channels can be well preserved. Experimental results on both the grayscale and color face images demonstrate the superiority of MRGR and MRGR-Q compared with several state-of-the-art methods. Licheng Liu, C. L. Philip Chen, Yaonan Wang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2023 | A Segmented Variable-Parameter ZNN for Dynamic Quadratic Minimization With Improved Convergence and RobustnessabstractAs a category of the recurrent neural network (RNN), zeroing neural network (ZNN) can effectively handle time-variant optimization issues. Compared with the fixed-parameter ZNN that needs to be adjusted frequently to achieve good performance, the conventional variable-parameter ZNN (VPZNN) does not require frequent adjustment, but its variable parameter will tend to infinity as time grows. Besides, the existing noise-tolerant ZNN model is not good enough to deal with time-varying noise. Therefore, a new-type segmented VPZNN (SVPZNN) for handling the dynamic quadratic minimization issue (DQMI) is presented in this work. Unlike the previous ZNNs, the SVPZNN includes an integral term and a nonlinear activation function, in addition to two specially constructed time-varying piecewise parameters. This structure keeps the time-varying parameters stable and makes the model have strong noise tolerance capability. Besides, theoretical analysis on SVPZNN is proposed to determine the upper bound of convergence time in the absence or presence of noise interference. Numerical simulations verify that SVPZNN has shorter convergence time and better robustness than existing ZNN models when handling DQMI. Lin Xiao 0002, Yongjun He 0001, Yaonan Wang 0001, Jianhua Dai 0003, Ran Wang 0001, Wensheng Tang |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2023 | Sharing Traffic Priorities via Cyber-Physical-Social Intelligence: A Lane-Free Autonomous Intersection Management Method in MetaverseabstractReplacing traffic signals with roadside vehicle-to-infrastructure systems in the era of connected and autonomous vehicles (CAVs) is promising. Managing CAVs in a signal-free intersection, known as autonomous intersection management (AIM), controls the driving behavior of each intersection-traverse CAV to maximize the throughput. Although AIM improves the gross throughput, the fairness of each individual vehicle in its right of way is not seriously considered. This study sets up an AIM system in the cyber–physical–social space to trade traverse priorities quantitatively and fairly. To that end, one needs an AIM method that is optimal and stable, otherwise no convincing trades of traverse priorities could be made. This study proposes a near-optimal lane-free AIM method based on numerical optimal control, wherein log-exp functions are deployed to convexify nondifferentiable collision-avoidance constraints. Besides that, a parameterized social force model (SFM) is proposed to provide a tunable initial guess for numerical optimal control. By tuning the urgency weights in SFM, one may get cooperative trajectories in different homotopy classes, which are further utilized to decide the amount of virtual currency to reward those CAVs who tend to share their traverse priorities. The overall method improves the traverse throughput with individual fairness respected. In experiencing this system, passengers learn how to behave with politeness when they drive manually. Experiments show the efficiency and robustness of the AIM method and also show the efficacy of the overall priority-sharing system. Bai Li 0002, Dongpu Cao, Hairong Dong 0001, Yaonan Wang 0001, Fei-Yue Wang 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 6 |
| 2023 | Adaptive Neural Design of Consensus Controllers for Nonlinear Multiagent Systems Under Switching TopologiesabstractExisting adaptive neural control methods for nonlinear multiagent systems (MASs) are only applicable under a fixed topology or are applicable under switching topologies but require some linear growth conditions on the nonlinear functions. Motivated by these limitations, a state-dependent adaptive neural design method is proposed in this article. Technically, our method is developed from a state-dependent Lyapunov function candidate, a switched control law, and a projection-based adaptation mechanism. To overcome the stability analysis difficulty caused by the new design of the Lyapunov function, a nonswitched compensation approach and a modified multiple Lyapunov functions method are proposed to derive a dwell-time condition, under which stability can be preserved. It is proved that in addition to stability, synchronization errors converge to a tunable residual around zero. Besides, the proposed scheme achieves the improvement of transient performance in terms of$L_{2}$norm and moreover, once there are no more topology switchings, asymptotic convergence of synchronization errors to a prescribed interval recovers automatically. Kaixin Lu, Zhi Liu 0001, Yaonan Wang 0001, C. L. Philip Chen, Yun Zhang 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 3 |
| 2023 | Consensus-Based Multipopulation Game Dynamics for Distributed Nash Equilibria Seeking and OptimizationabstractIn this article, a consensus-based multipopulation game dynamics approach is proposed for distributed Nash equilibria seeking and optimization. The approach is fundamentally different from existing population dynamics from that: 1) the underlying communication network underlying the population game dynamics could be arbitrary undirected connected graph and more importantly 2) the proposed approach works in a distributed manner even when the objective functions of players are strongly coupled with each other. The proposed approach greatly extends the applicability of population game dynamics in distributed optimization and learning problems. A distributed constrained optimization problem and a traffic routing problem are utilized to illustrate the feasibility of the proposed multipopulation game dynamics approach. Shaolin Tan, Zhihong Fang, Yaonan Wang 0001, Jinhu Lü 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 3 |
| 2023 | Erratum to "Distributed Population Dynamics for Searching Generalized Nash Equilibria of Population Games With Graphical Strategy Interactions"abstractIn[1], the affiliation for Athanasios V. Vasilakos should be as follows: Shaolin Tan, Yaonan Wang 0001, Athanasios V. Vasilakos |
IEEE Trans. Syst. Man Cybern. Syst. | 2 |
| 2023 | A Composite Control Framework of Safety Satisfaction and Uncertainties Compensation for Constrained Time-Varying Nonlinear MIMO SystemsabstractIn this article, we propose a composite control framework for time-varying nonlinear multiple-input–multiple-output (MIMO) systems with safety constraints and unknown dynamics. The control framework combines control barrier functions (CBFs) and high-order CBFs (HoCBFs) with an extended state observer (ESO) to handle arbitrary relative-degree constraints in the presence of system uncertainties and output measurements only. Then, the ESO-CBF/HoCBF-based safety control schemes are obtained by solving quadratic programs (QPs) to ensure the safety of closed-loop control systems. Consequently, the safety satisfaction and uncertainties compensation objectives can be achieved simultaneously. Finally, simulations and experiments are conducted to verify the effectiveness of the proposed safety control schemes. Haijing Wang, Jinzhu Peng, Fangfang Zhang 0004, Yaonan Wang 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 4 |
| 2023 | Adaptive Neural Prescribed-Time Control of Switched Nonlinear Systems With Mode-Dependent Average Dwell TimeabstractMost current adaptive neural control strategies for switched nonlinear systems, both finite-time and fixed-time, are limited by a conservatively estimated settling time. Besides, the convergence accuracy of these methods is only bounded but unknown and uncertain. This study proposes a neural adaptive prescribed-time control method to solve such a problem. Specifically, a critical design step is that the practical prescribed-time control problem is converted into a practical stabilization problem by developing a new singularity-avoidance error-dependent scalar transformation. Guided by this idea, an adaptive neural prescribed-time controller is constructed, ensuring prescribed transient behavior and all tracking errors to achieve preset accuracy within the prescribed time simultaneously. Furthermore, by utilizing extended multiple Lyapunov functions, a new mode-dependent average dwell time condition is derived to ensure that all signals in the controlled system remain bounded. Finally, simulations demonstrate the feasibility of the developed scheme. Danping Zeng, Zhi Liu 0001, C. L. Philip Chen, Yaonan Wang 0001, Yun Zhang 0001, Zongze Wu 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 4 |
| 2023 | Asymptotic Stabilization Control of Fractional-Order Memristor-Based Neural Networks System via Combining Vector Lyapunov Function With M-MatrixabstractThis article examines a new measure of combining the vector Lyapunov function with${M}$-matrix for settling the asymptotic stabilization control of fractional-order memristor-based neural networks system (FOMBNNS) has large delays in various dimensional forms. Some new stability and stabilization criteria are deduced. First, the vector Lyapunov function and${M}$-matrix are imported for investigating stabilization control for the above system. Then, we solve the problem for a special type of situation that the activation functions no longer consider Lipschitz parameters via the new method. Finally, four numerical examples from different kinds of situations are simulated for expounding the validity of the novel asymptotic stability and stabilization criteria. Compared with the methods mentioned in the current references, the proposed asymptotic stability and stabilization criteria in this article have strong generality and universality. They can be applied not only to the most common feedback control, accordingly, the feedback control law based on which they are designed but also to all fractional-order parameters from 0 to 1. In addition, the new method has lower conservativeness and fewer constraints. Moreover, the new stability and stabilization criteria can also overcome the difficulty in dealing with the above system owning large delays. Zhe Zhang 0022, Yaonan Wang 0001, Jing Zhang 0014, Hui Zhang 0023, Zhaoyang Ai, Kan Liu 0002 |
IEEE Trans. Syst. Man Cybern. Syst. | 2 |
| 2022 | Learning from Pixel-Level Noisy Label : A New Perspective for Light Field Saliency DetectionabstractSaliency detection with light field images is becoming attractive given the abundant cues available, however, this comes at the expense of large-scale pixel level annotated data which is expensive to generate. In this paper, we propose to learn light field saliency from pixel-level noisy labels obtained from unsupervised hand crafted featured-based saliency methods. Given this goal, a natural question is: can we efficiently incorporate the relationships among light field cues while identifying clean labels in a unified framework? We address this question by formulating the learning as a joint optimization of intra light field features fusion stream and inter scenes correlation stream to generate the predictions. Specially, we first introduce a pixel forgetting guided fusion module to mutually enhance the light field features and exploit pixel consistency across iterations to identify noisy pixels. Next, we introduce a cross scene noise penalty loss for better reflecting latent structures of training data and enabling the learning to be invariant to noise. Extensive experiments on multiple benchmark datasets demonstrate the superiority of our framework showing that it learns saliency prediction comparable to state-of-the-art fully supervised light field saliency methods. Our code is available at h t tps://github.com/ OLobbCode/NoiseLF. Mingtao Feng, Kendong Liu, Liang Zhang 0010, Hongshan Yu, Yaonan Wang 0001, Ajmal Mian |
CVPR | 5 |
| 2022 | Negative Stiffness Analysis and Regulation of In-Hand Manipulation with Underactuated Compliant HandsabstractThis paper addresses the generation mechanism and avoidance method of negative stiffness during in-Hand manipulation with underactuated compliant hands. Firstly, a planar hand with two three-jointed fingers manipulating a rectangular is set, and a quasi-static underactuated operation model is established. Secondly, based on this simulation model, we investigated the stiffness evolution during in-hand manipulation, and analyze the influence factors of system stiffness. Finally, a stiffness regulation method is developed to avoid negative stiffness during in-hand manipulation. The method is validated by simulation. The research results are beneficial to improve the performance of underactuated in-hand manipulation. Wenrui Chen, Qiang Diao, Yaonan Wang 0001, Cuo Yan, Zhiyong Li 0001 |
ICRA | 3 |
| 2022 | ICK-Track: A Category-Level 6-DoF Pose Tracker Using Inter-Frame Consistent Keypoints for Aerial ManipulationabstractRobots that are supposed to interact with or manipulate objects in the world must be able to track the poses of objects in their sensor data. Thus, Detecting and tracking the 6-DoF poses of targeted objects is important for aerial manipulation and is still in the early stage due to the high dynamics and limited onboard capacity of such systems. In this paper, we propose ICK-Track, a novel method for onboard category-level object 6-DoF pose tracking that can be applied to aerial manipulation without using any pre-defined object CAD models. It first utilizes a semi-supervised video segmentation to detect objects in the eye-in-hand RGB-D camera stream to segment the 3D points of objects. Then, canonical keypoints are extracted using iterative farthest point sampling. We propose a novel inter-frame consistent keypoints generation network to generate the corresponding keypoint pairs, which are used together with ICP to estimate the pose changes of objects for tracking. Experimental results show that our method is more robust to viewpoint changes and runs faster than the state-of-the-art methods on category-level pose tracking. We further test our proposed method on a real aerial manipulator. A demo video showing the use of our method on a real aerial manipulator and the implementation of our method are available at: https://github.com/S-JingTao/ICK-Track. Yaonan Wang 0001, Mingtao Feng, Danwei Wang, Jiawen Zhao, Cyrill Stachniss, Xieyuanli Chen |
IROS | 2 |
| 2022 | Enhancement of Target Feature Regions and Intention-Driven Visual Attention Selection in Traffic ScenesabstractSelective attention to specific areas and specific targets according to driving intention is of great significance for autonomous vehicles to efficiently obtain external environment information. In order to achieve efficient environmental perception, we propose an intention-driven visual attention selection model by simulating human active perception of the external environment. Meanwhile, in order to improve the integrity of the target category attention heatmap, a deep network training method with feature region enhancement is proposed. In this paper, FIMF Score-CAM which can fast integrate multiple features of local space is proposed. It generates intention-related target attention map by weighting the feature map extracted by forward convolution calculation, and combines spatial attention and feature attention to improve the ability of target category location. At the same time, the network is forced to pay more attention to the more comprehensive target-related region by using the guided random erasing in training process, which overcomes the deficiency that the model only pays attention to the most discriminative feature region, and achieves the purpose of feature region enhancement. Experiments on KITTI dataset show that the positioning integrity and accuracy of our model are significantly improved compared with other top-down attention models. Jing Li 0174, Dongbo Zhang 0003, Bumin Meng, Yaonan Wang 0001 |
IV | 6 |
| 2022 | Correlation filter tracking algorithm based on spatial-temporal regularization and context awareness
Xuedong Wu, Yaonan Wang 0001, Siming Tang, Mengquan Liang, Baiheng Cao |
Appl. Intell. | 4 |
| 2022 | Accurate RGB-D SLAM in dynamic environments based on dynamic visual feature removal
Chenxin Liu, Jiahu Qin, Shuai Wang 0018, Yaonan Wang 0001 |
Sci. China Inf. Sci. | 5 |
| 2022 | Review on the COVID-19 pandemic prevention and control system based on AI
Junfei Yi, Hui Zhang 0023, Jianxu Mao, Yurong Chen 0003, Hang Zhong, Yaonan Wang 0001 |
Eng. Appl. Artif. Intell. | 6 |
| 2022 | A payoff-based learning approach for Nash equilibrium seeking in continuous potential games
Shaolin Tan, Yaonan Wang 0001 |
Neurocomputing | 2 |
| 2022 | Self-supervised part segmentation via motion imitation
Qiaokang Liang, Kunlin Zou, Wei Sun 0028, Yaonan Wang 0001 |
Image Vis. Comput. | 6 |
| 2022 | Control-oriented modeling and optimization for the temperature and airflow management in an air-cooled data-center
Qiu Fang, Jiakang Zhou, Yaonan Wang 0001 |
Neural Comput. Appl. | 4 |
| 2022 | A Deep Local Patch Matching Network for Cell Tracking in Microscopy Image Sequences Without RegistrationabstractCell tracking is critical for the modeling of plant cell growth patterns. A local graph matching algorithm is proposed to track cells by exploiting the tight spatial topology of cells. However, the local graph matching approach lacks robustness in the unregistered images because the feature descriptors are handcrafted. In this paper, we propose a Deep Local Patch Matching Network (DLPM-Net) to track cells robustly, by exploiting local patches' deep similarity information and cells' spatial-temporal contextual information. Furthermore, to reduce the time consumption during the matching process and enhance tracking accuracy, we take two steps to realize the tracking of non-division cells and the detection of cell divisions. In the first step, the DLPM-Net is employed to match the non-division cells by exploiting the cell pair candidates' local patch contextual information, then the non-matched cells are recorded as the cell division candidates. In the second step, the DLPM-Net is used to detect cell divisions from these non-matched cells, by exploiting the local patch contextual similarity between the mother cell's local patch and daughter cells' local patch. Compared with the existing local graph matching method, the experimental results show that the proposed method gains 29.1% improvement in the tracking accuracy. Yulian Xie, Min Liu 0008, Shirui Zhou, Yaonan Wang 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 4 |
| 2022 | Spatial Decomposition-Based Fault Detection Framework for Parabolic-Distributed Parameter ProcessesabstractFault detection for distributed parameter processes (heat processes, fluid processes, etc.) is vital for safe and efficient operation. On one hand, the existing data-driven methods neglect the evolution dynamics of the processes and cannot guarantee that they work for highly dynamic or transient processes; on the other hand, model-based methods reported so far are mostly based on the backstepping technique, which does not possess enough redundancy for fault detection since only the boundary measurement is considered. Motivated by these considerations, we intend to investigate the robust fault detection problem for distributed parameter processes in a model-based perspective covering both boundary and in-domain measurement cases. A real-time fault detection filter (FDF) is presented, which gets rid of a large amount of data collection and offline training procedures. Rigorous theoretic analysis is presented for guiding the parameters selection and threshold computation. A time-varying threshold is designed such that the false alarm in the transient stage can be avoided. Successful application results on a hot strip mill cooling system demonstrate the potential for real industrial applications. Yun Feng 0001, Yaonan Wang 0001, Bing-Chuan Wang, Han-Xiong Li |
IEEE Trans. Cybern. | 2 |
| 2022 | Resilient Adaptive Neural Control for Uncertain Nonlinear Systems With Infinite Number of Time-Varying Actuator FailuresabstractExisting studies on adaptive fault-tolerant control for uncertain nonlinear systems with actuator failures are restricted to a common result that only system stability is established. Such a result of not being asymptotically stable is a tradeoff paid for reducing the number of online learning parameters. In this article, we aim to obviate such restrictions and improve the bounded error control to asymptotic control. Toward this end, a resilient adaptive neural control scheme is newly proposed based on a new design of the Lyapunov function candidates, a projection-associated tuning functions method, and an alternative class of smooth functions. It is proved that the system stability is guaranteed for the case of an infinite number of failures and when the number of failures is finite, asymptotic tracking performance can be automatically recovered, and besides, an explicit bound for the tracking error in terms of$L_{2}$norm is established. Illustrative examples demonstrate the methods developed. Kaixin Lu, Zhi Liu 0001, Yaonan Wang 0001, C. L. Philip Chen |
IEEE Trans. Cybern. | 3 |
| 2022 | Inverse Optimal Design of Direct Adaptive Fuzzy Controllers for Uncertain Nonlinear SystemsabstractOptimized performance obtained from existing adaptive fuzzy optimal control methods comes at the cost of a intricate design procedure and a heavy computation of online parameter learning, and it is an under-explored problem on how to remove such a restriction. In this article, we tackle this problem and ensure the optimized performance using only one adaptive parameter. To this end, a direct adaptive fuzzy inverse approach is first proposed to design a switching-type inverse optimal controller and a one-parameter learning mechanism. It is proved that the proposed approach ensures the input-to-state stability of the control system and besides, the inverse optimality in regard to a meaningful cost functional is achieved. Illustrative examples verify the approach developed. Kaixin Lu, Zhi Liu 0001, C. L. Philip Chen, Yaonan Wang 0001, Yun Zhang 0001 |
IEEE Trans. Fuzzy Syst. | 4 |
| 2022 | Design and Analysis of a Noise-Resistant ZNN Model for Settling Time-Variant Linear Matrix Inequality in Predefined-TimeabstractAiming at the efficient online solution of the time-variant linear matrix inequality (LMI) under nonideal conditions (e.g., noise pollution), a predefined-time convergent and integral-enhanced zeroing neural network (PCIE-ZNN) model is built for the first time in this article. Compared with existing zeroing neural network (ZNN) models for settling the time-variant LMI, the PCIE-ZNN model proposed in this article is proved to have better convergence and stronger robustness even in the presence of noise interference through strict mathematical analysis and detailed numerical simulations. Specifically, the stability, predefined-time convergence, and robustness of the PCIE-ZNN model are guaranteed in theory. Then, numerical simulation cases fully compare the results of the proposed PCIE-ZNN model and the existing ZNN models for the time-variant LMI, which demonstrates the correctness of theoretical proof and the superiority of the PCIE-ZNN model in settling the time-variant LMI under various noise pollution. In addition, through comparative experiments of three sets of design parameters, the convergence speed of the PCIE-ZNN model can be further accelerated by selecting proper parameters. Lin Xiao 0002, Wentong Song, Lei Jia 0001, Jiayue Sun, Yaonan Wang 0001 |
IEEE Trans. Ind. Informatics | 6 |
| 2022 | Low-Complexity Prescribed Performance Control for Unmanned Aerial Manipulator Robot System Under Model Uncertainty and Unknown DisturbancesabstractThis article presents a trajectory tracking control method for the unmanned aerial manipulator robot system (UAMRS) under model uncertainty and unknown disturbances. More specifically, a low-complexity prescribed performance controller is proposed to effectively reduce the design complexity and achieve the prescribed transient and steady-state performance. First, the dynamics model of the UAMRS is analyzed and modeled, where the unmeasured internal interaction generated by the coupling effect and the random environmental disturbances are considered simultaneously. Then, utilizing the property of prescribed performance, the UAMRS with model uncertainty and external disturbances can guarantee preferable trajectory tracking responses, where the nonlinear disturbance observer is used to estimate and compensate uncertainties and external disturbances. Moreover, the proposed controller defined by simple expressions does not require accurate knowledge of the UAMRS, which is of low complexity and can effectively reduce the amount of calculation. The stability of the proposed controller is analyzed. Finally, the performances of the proposed scheme are demonstrated by the numerical simulation comparisons and real-world experiments, where a quadrotor with a 3-DOF onboard active manipulator is adopted in outdoor experimental validations. Jiacheng Liang, Ningbin Lai, Bingwei He, Zhiqiang Miao, Yaonan Wang 0001 |
IEEE Trans. Ind. Informatics | 6 |
| 2022 | Low-Complexity Control for Vision-Based Landing of Quadrotor UAV on Unknown Moving PlatformabstractThis article addresses the vision-based landing problem of a low-cost quadrotor on an unknown moving platform. A robust landing controller is developed, which consists of the design of the low-complexity outer-loop controller and the geometric inner-loop attitude controller. First, an error transformation based on prescribed performance is designed to guarantee the landing behaviors and deal with the intermediate control signal of backstepping approaches, resulting in a low-complexity position-based visual servoing (PBVS) design. In addition, the proposed PBVS controller exhibits strong robustness against an uncertain relative dynamic system due to no incorporation of any prior knowledge of the moving platform. Next, a modified geometric attitude controller is presented by characterizing the geometric properties of rotation matrices intrinsically. Finally, the stability analysis is presented using Lyapunov stability theory, and the effectiveness of the proposed control strategy is demonstrated through numerical simulations and experiments. Jie Lin 0009, Yaonan Wang 0001, Zhiqiang Miao, Hang Zhong, Rafael Fierro |
IEEE Trans. Ind. Informatics | 2 |
| 2022 | Low-Complexity Leader-Following Formation Control of Mobile Robots Using Only FOV-Constrained Visual FeedbackabstractThis article aims to solve the problem of formation control of mobile robots based on image and provide a low-cost as well as ease-of-implementation solution for mobile robots relying merely on a monocular camera under field-of-view (FOV) constraints. A low-complexity image-based visual servo controller is proposed, which can achieve the desired relative position on the image plane and solve the FOV constraints without the feature depth and leader’s velocities information. To facilitate the control design, a state transformation is first performed to decouple the visual motion kinematics. Then, an error transformation is introduced to handle the FOV constraints, and performance specifications are incorporated in the error transformation to achieve the predefined control performance. Finally, a simple static controller is derived using only information from images, and the stability of the uncertain system with unknown control direction/coefficients under the given performance control condition is analyzed. The effectiveness and performance of the proposed visual servoing controller can be illustrated using both simulations and experiments. Zhiqiang Miao, Hang Zhong, Yaonan Wang 0001, Hui Zhang 0023, Haoran Tan, Rafael Fierro |
IEEE Trans. Ind. Informatics | 3 |
| 2022 | Deep Stereo Matching With Hysteresis Attention and Supervised Cost Volume ConstructionabstractStereo matching disparity prediction for rectified image pairs is of great importance to many vision tasks such as depth sensing and autonomous driving. Previous work on the end-to-end unary trained networks follows the pipeline of feature extraction, cost volume construction, matching cost aggregation, and disparity regression. In this paper, we propose a deep neural network architecture for stereo matching aiming at improving the first and second stages of the matching pipeline. Specifically, we show a network design inspired by hysteresis comparator in the circuit as our attention mechanism. Our attention module is multiple-block and generates an attentive feature directly from the input. The cost volume is constructed in a supervised way. We try to use data-driven to find a good balance between informativeness and compactness of extracted feature maps. The proposed approach is evaluated on several benchmark datasets. Experimental results demonstrate that our method outperforms previous methods on SceneFlow, KITTI 2012, and KITTI 2015 datasets. Kai Zeng 0010, Yaonan Wang 0001, Jianxu Mao, Caiping Liu, Weixing Peng, Yin Yang 0004 |
IEEE Trans. Image Process. | 2 |
| 2022 | DeepRayburst for Automatic Shape Analysis of Tree-Like Structures in Biomedical ImagesabstractPrecise quantification of tree-like structures from biomedical images, such as neuronal shape reconstruction and retinal blood vessel caliber estimation, is increasingly important in understanding normal function and pathologic processes in biology. Some handcrafted methods have been proposed for this purpose in recent years. However, they are designed only for a specific application. In this paper, we propose a shape analysis algorithm, DeepRayburst, that can be applied to many different applications based on a Multi-Feature Rayburst Sampling (MFRS) and a Dual Channel Temporal Convolutional Network (DC-TCN). Specifically, we first generate a Rayburst Sampling (RS) core containing a set of multidirectional rays. Then the MFRS is designed by extending each ray of the RS to multiple parallel rays which extract a set of feature sequences. A Gaussian kernel is then used to fuse these feature sequences and outputs one feature sequence. Furthermore, we design a DC-TCN to make the rays terminate on the surface of tree-like structures according to the fused feature sequence. Finally, by analyzing the distribution patterns of the terminated rays, the algorithm can serve multiple shape analysis applications of tree-like structures. Experiments on three different applications, including soma shape reconstruction, neuronal shape reconstruction, and vessel caliber estimation, confirm that the proposed method outperforms other state-of-the-art shape analysis methods, which demonstrate its flexibility and robustness. Weixun Chen, Min Liu 0008, Yaonan Wang 0001, Erik Meijering |
IEEE J. Biomed. Health Informatics | 4 |
| 2022 | An O-Shape Neural Network With Attention Modules to Detect Junctions in Biomedical Images Without SegmentationabstractJunction plays an important role in biomedical research such as retinal biometric identification, retinal image registration, eye-related disease diagnosis and neuron reconstruction. However, junction detection in original biomedical images is extremely challenging. For example, retinal images contain many tiny blood vessels with complicated structures and low contrast, which makes it challenging to detect junctions. In this paper, we propose an O-shape Network architecture with Attention modules (Attention O-Net), which includes Junction Detection Branch (JDB) and Local Enhancement Branch (LEB) to detect junctions in biomedical images without segmentation. In JDB, the heatmap indicating the probabilities of junctions is estimated and followed by choosing the positions with the local highest value as the junctions, whereas it is challenging to detect junctions when the images contain weak filament signals. Therefore, LEB is constructed to enhance the thin branch foreground and make the network pay more attention to the regions with low contrast, which is helpful to alleviate the imbalance of the foreground between thin and thick branches and to detect the junctions of the thin branch. Furthermore, attention modules are utilized to introduce the feature maps of LEB to JDB, which can establish a complementary relationship and further integrate local features and contextual information between these two branches. The proposed method achieves the highest average F1-scores of 0.82, 0.73 and 0.94 in two retinal datasets and one neuron dataset, respectively. The experimental results confirm that Attention O-Net outperforms other state-of-the-art detection methods, and is helpful for retinal biometric identification. Min Liu 0008, Fuhao Yu, Tieyong Zeng, Yaonan Wang 0001 |
IEEE J. Biomed. Health Informatics | 5 |
| 2022 | 3D Gradient Reconstruction-Based Path Planning Method for Autonomous Vehicle With Enhanced Roll StabilityabstractPath planning has received more and more attention due to its indispensability in autonomous vehicle (AV). Generally speaking, the stability of AV is not fully considered in path planning, therefore, the planned path may be detrimental for maintaining the stability. And this problem is even more acute for roll stability in complex 3D environment, such as off-road terrain environment. To enhance the roll stability on off-road terrain, a 3D gradient reconstruction-based path planning method (3DGRB-PPM) is presented in this work. The 3DGRB-PPM can keep the roll angle at a low level even in complex 3D environment, thus greatly enhancing the roll stability. The 3DGRB-PPM includes two parts, a gradient reconstruction unit and an adaptive fusion unit. In order to ensure the roll stability of AV in path planning, the gradient reconstruction unit is designed by constructing two potential fields, a joint potential field of gradient and roll angle for enhancing the roll stability and a 3D artificial potential field for reaching the destination. In order to coordinate these two potential fields, an adaptive fusion unit is designed by fuzzy inference rule. The simulation is implemented on the Matlab-Carsim co-simulation platform, and the simulation results show that the path planned by 3DGRB-PPM has good performance with enhanced roll stability. Zhixian Liu, Xiaofang Yuan, Guoming Huang, Weihua Tan, Yaonan Wang 0001 |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2022 | A Model of Extraction of Rail's Vertical Corrugation Based on Flexible Virtual RulerabstractRail corrugation (RC) is one of the most important indicator to evaluate the quality of rail, which is used to describe the irregularity of rail surface, also known as rail’s vertical flatness. However, the definition of RC is still an empirical description for low-speed measurement devices. In this paper, the process of RC measurement is divided into two steps, sampling and extraction, which helps the users to understand more clearly. Then a new mathematic model is proposed to make the process of RC’s extraction be an executable operation for machine calculation. The model adopts a new concept of flexible virtual ruler to perform sliding filtering on an overall rail and extract the instantaneous RC with a successive approximation algorithm according to user’s requirements and national standards. The proposed FVR model can not only be fully compatible with the traditional extraction method of RC, but also provide a new idea for evaluating RC’s quality. Comparing with the current popular methods, the proposed model gives a meaningful strategy for rail maintenance with a complete mathematic description having more degrees of freedom. Experiment results demonstrate its validity and reliability for both indoor simulation and actual outdoor experiments. Ziji Ma, Kehuang Xu, Xun Shao, Mianxiong Dong, Yaonan Wang 0001 |
IEEE Trans. Intell. Transp. Syst. | 6 |
| 2022 | Deep Progressive Fusion Stereo NetworkabstractStereo matching depth estimation for rectified image pairs is of great importance to many compute vision tasks, specifically in autonomous driving. With the flourishing of convolution neural networks, responsible depth estimation of stereo matching with artificial intelligence is the most severe challenge for autonomous driving in recent years. Previous research on end-to-end trainable stereo matching networks has usually used cascading convolution blocks with down-sampling or pooling operations to extract the unary features required for matching cost construction. Such approaches lack a reconstruction stage for increasing feature map pixel-wise alignment and strength, factors which play an important role in representing the similarity between stereo image pairs. To address this issue, in this paper, we propose the progressive fusion stereo matching network (PFSM-Net). We exploit an encoder-decoder feature extraction network architecture for multi-stage and -scale dynamic feature extraction. Moreover, we propose a group-wise concatenation method to construct the cost volume, which provides a more efficient cost volume for cost aggregation. Furthermore, we propose the use of multi-scale cost aggregation networks with a progressive fusion strategy. The aggregated cost volume is progressively fused with the multi-stage and -scale cost volume as the size of the cost volume increases. Multi-stage and -scale outputs are supervised with and learned in a coarse-to-fine manner. Experimental results demonstrate that our method outperforms previous methods on the SceneFlow, KITTI 2012, and KITTI 2015 datasets. Kai Zeng 0010, Yaonan Wang 0001, Qing Zhu 0003, Jianxu Mao, Hui Zhang 0023 |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2022 | MRSDI-CNN: Multi-Model Rail Surface Defect Inspection System Based on Convolutional Neural NetworksabstractDefects on rail surfaces, which have become critical problems, need to be detected and removed as quickly as possible to ensure the fast, safe, and stable operation of trains. At present, although many solutions have been proposed to address these problems, the comprehensiveness, rapidity, and accuracy of defect detection remain unsatisfactory. This study aims to resolve these existing problems and accordingly proposes a multi-model rail surface defect detection system based on convolutional neural networks (MRSDI-CNN) from the standpoint of studying the squat on the rail surface. The convolutional neural networks utilized include the improved Single Shot MultiBox Detector (SSD) and You Only Look Once version 3(YOLOv3)—two types of one-stage networks. We expounded and analyzed the performance of the convolutional neural networks as well as their applicability to rail surface defect detection. We used a diverse range of rail defect sizes to improve the detection performance of the two deep learning networks, following which they could identify three types of squats in parallel with improved accuracy and without reduction of the detection speed. The experimental results confirm the effectiveness and superiority of the proposed method over those of previous studies. Hui Zhang 0023, Yanan Song, Yurong Chen 0003, Hang Zhong, Li Liu 0060, Yaonan Wang 0001, Akilan Thangarajah, Q. M. Jonathan Wu |
IEEE Trans. Intell. Transp. Syst. | 6 |
| 2022 | Deep-Learning-Based Automated Neuron Reconstruction From 3D Microscopy Images Using Synthetic Training ImagesabstractDigital reconstruction of neuronal structures from 3D microscopy images is critical for the quantitative investigation of brain circuits and functions. It is a challenging task that would greatly benefit from automatic neuron reconstruction methods. In this paper, we propose a novel method called SPE-DNR that combines spherical-patches extraction (SPE) and deep-learning for neuron reconstruction (DNR). Based on 2D Convolutional Neural Networks (CNNs) and the intensity distribution features extracted by SPE, it determines the tracing directions and classifies voxels into foreground or background. This way, starting from a set of seed points, it automatically traces the neurite centerlines and determines when to stop tracing. To avoid errors caused by imperfect manual reconstructions, we develop an image synthesizing scheme to generate synthetic training images with exact reconstructions. This scheme simulates 3D microscopy imaging conditions as well as structural defects, such as gaps and abrupt radii changes, to improve the visual realism of the synthetic images. To demonstrate the applicability and generalizability of SPE-DNR, we test it on 67 real 3D neuron microscopy images from three datasets. The experimental results show that the proposed SPE-DNR method is robust and competitive compared with other state-of-the-art neuron reconstruction methods. Weixun Chen, Min Liu 0008, Miroslav Radojevic, Yaonan Wang 0001, Erik Meijering |
IEEE Trans. Medical Imaging | 5 |
| 2022 | A 3D Tubular Flux Model for Centerline Extraction in Neuron Volumetric ImagesabstractDigital morphology reconstruction from neuron volumetric images is essential for computational neuroscience. The centerline of the axonal and dendritic tree provides an effective shape representation and serves as a basis for further neuron reconstruction. However, it is still a challenge to directly extract the accurate centerline from the complex neuron structure with poor image quality. In this paper, we propose a neuron centerline extraction method based on a 3D tubular flux model via a two-stage CNN framework. In the first stage, a 3D CNN is used to learn the latent neuron structure features, namely flux features, from neuron images. In the second stage, a light-weight U-Net takes the learned flux features as input to extract the centerline with a spatial weighted average strategy to constrain the multi-voxel width response. Specifically, the labels of flux features in the first stage are generated by the 3D tubular model which calculates the geometric representations of the flux between each voxel in the tubular region and the nearest point on the centerline ground truth. Compared with self-learned features by networks, flux features, as a kind of prior knowledge, explicitly take advantage of the contextual distance and direction distribution information around the centerline, which is beneficial for the precise centerline extraction. Experiments on two challenging datasets demonstrate that the proposed method outperforms other state-of-the-art methods by 18% and 35.1% in F1-measurement and average distance scores at the most, and the extracted centerline is helpful to improve the neuron reconstruction performance. Min Liu 0008, Yaonan Wang 0001, Jiawang Fan, Erik Meijering |
IEEE Trans. Medical Imaging | 3 |
| 2022 | Structure-Guided Segmentation for 3D Neuron ReconstructionabstractDigital reconstruction of neuronal morphologies in 3D microscopy images is critical in the field of neuroscience. However, most existing automatic tracing algorithms cannot obtain accurate neuron reconstruction when processing 3D neuron images contaminated by strong background noises or containing weak filament signals. In this paper, we present a 3D neuron segmentation network named Structure-Guided Segmentation Network (SGSNet) to enhance weak neuronal structures and remove background noises. The network contains a shared encoding path but utilizes two decoding paths called Main Segmentation Branch (MSB) and Structure-Detection Branch (SDB), respectively. MSB is trained on binary labels to acquire the 3D neuron image segmentation maps. However, the segmentation results in challenging datasets often contain structural errors, such as discontinued segments of the weak-signal neuronal structures and missing filaments due to low signal-to-noise ratio (SNR). Therefore, SDB is presented to detect the neuronal structures by regressing neuron distance transform maps. Furthermore, a Structure Attention Module (SAM) is designed to integrate the multi-scale feature maps of the two decoding paths, and provide contextual guidance of structural features from SDB to MSB to improve the final segmentation performance. In the experiments, we evaluate our model in two challenging 3D neuron image datasets, the BigNeuron dataset and the Extended Whole Mouse Brain Sub-image (EWMBS) dataset. When using different tracing methods on the segmented images produced by our method rather than other state-of-the-art segmentation methods, the distance scores gain 42.48% and 35.83% improvement in the BigNeuron dataset and 37.75% and 23.13% in the EWMBS dataset. Bo Yang 0065, Min Liu 0008, Yaonan Wang 0001, Erik Meijering |
IEEE Trans. Medical Imaging | 3 |
| 2022 | Composite-Learning-Based Adaptive Neural Control for Dual-Arm Robots With Relative MotionabstractThis article presents an adaptive control method for dual-arm robot systems to perform bimanual tasks under modeling uncertainties. Different from the traditional symmetric bimanual robot control, we study the dual-arm robot control with relative motions between robotic arms and a grasped object. The robot system is first divided into two subsystems: a settled manipulator system and a tool-used manipulator system. Then, a command filtered control technique is developed for trajectory tracking and contact force control. In addition, to deal with the inevitable dynamic uncertainties, a radial basis function neural network (RBFNN) is employed for the robot, with a novel composite learning law to update the NN weights. The composite learning is mainly based on an integration of the historic data of NN regression such that information of the estimate error can be utilized to improve the convergence. Moreover, a partial persistent excitation condition is employed to ensure estimation convergence. The stability analysis is performed by using the Lyapunov theorem. Numerical simulation results demonstrate the validity of the proposed control and learning algorithm. Yiming Jiang 0001, Yaonan Wang 0001, Zhiqiang Miao, Jing Na, Zhijia Zhao 0002, Chenguang Yang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2022 | Bio-Inspired Dynamic Collective Choice in Large-Population Systems: A Robust Mean-Field Game PerspectiveabstractInspired by the collective decision making in biological systems, such as honeybee swarm searching for a new colony, we study a dynamic collective choice problem for large-population systems with the purpose of realizing certain advantageous features observed in biology. This problem focuses on the situation where a large number of heterogeneous agents subject to adversarial disturbances move from initial positions toward one of the destinations in a finite time while trying to remain close to the average trajectory of all agents. To overcome the complexity of this problem resulting from the large population and the heterogeneity of agents, and also to enforce some specific choices by individuals, we formulate the problem under consideration as a robust mean-field game with non-convex and non-smooth cost functions. Through Nash equivalence principle, we first deal with a single-player$H_{\infty }$tracking problem by taking the population behavior as a fixed trajectory, and then establish a mean-field system to estimate the population behavior. Optimal control strategies and worst disturbances, independent of the population size, are designed, which give a way to realize the collective decision-making behavior emerged in biological systems. We further prove that the designed strategies constitute$\epsilon _{N}$-Nash equilibrium, where$\epsilon _{N}$goes toward zero as the number of agents increases to infinity. The effectiveness of the proposed results are illustrated through two simulation examples. Man Li 0002, Jiahu Qin, Yaonan Wang 0001, Yu Kang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2022 | Noise Robust Face Hallucination Based on Smooth Correntropy RepresentationabstractFace hallucination technologies have been widely developed during the past decades, among which the sparse manifold learning (SML)-based approaches have become the popular ones and achieved promising performance. However, these SML methods always failed in handling noisy images due to the least-square regression (LSR) they used for error approximation. To this end, we propose, in this article, a smooth correntropy representation (SCR) model for noisy face hallucination. In SCR, the correntropy regularization and smooth constraint are combined into one unified framework to improve the resolution of noisy face images. Specifically, we introduce the correntropy induced metric (CIM) rather than the LSR to regularize the encoding errors, which admits the proposed method robust to noise with uncertain distributions. Besides, the fused LASSO penalty is added into the feature space to ensure similar training samples holding similar representation coefficients. This encourages the SCR not only robust to noise but also can well exploit the inherent typological structure of patch manifold, resulting in more accurate representations in noise environment. Comparison experiments against several state-of-the-art methods demonstrate the superiority of SCR in super-resolving noisy low-resolution (LR) face images. Licheng Liu, Qiying Feng, C. L. Philip Chen, Yaonan Wang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2022 | Distributed Group Coordination of Multiagent Systems in Cloud Computing Systems Using a Model-Free Adaptive Predictive Control StrategyabstractThis article studies the group coordinated control problem for distributed nonlinear multiagent systems (MASs) with unknown dynamics. Cloud computing systems are employed to divide agents into groups and establish networked distributed multigroup-agent systems (ND-MGASs). To achieve the coordination of all agents and actively compensate for communication network delays, a novel networked model-free adaptive predictive control (NMFAPC) strategy combining networked predictive control theory with model-free adaptive control method is proposed. In the NMFAPC strategy, each nonlinear agent is described as a time-varying data model, which only relies on the system measurement data for adaptive learning. To analyze the system performance, a simultaneous analysis method for stability and consensus of ND-MGASs is presented. Finally, the effectiveness and practicability of the proposed NMFAPC strategy are verified by numerical simulations and experimental examples. The achievement also provides a solution for the coordination of large-scale nonlinear MASs. Haoran Tan, Yaonan Wang 0001, Min Wu 0002, Zhiwu Huang, Zhiqiang Miao |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2022 | Performance Analysis and Applications of Finite-Time ZNN Models With Constant/Fuzzy Parameters for TVQPEIabstractBased on extensive applications of the time-variant quadratic programming with equality and inequality constraints (TVQPEI) problem and the effectiveness of the zeroing neural network (ZNN) to address time-variant problems, this article proposes a novel finite-time ZNN (FT-ZNN) model with a combined activation function, aimed at providing a superior efficient neurodynamic method to solve the TVQPEI problem. The remarkable properties of the FT-ZNN model are faster finite-time convergence and preferable robustness, which are analyzed in detail, where in the case of the robustness discussion, two kinds of noises (i.e., bounded constant noise and bounded time-variant noise) are taken into account. Moreover, the proposed several theorems all compute the convergent time of the nondisturbed FT-ZNN model and the disturbed FT-ZNN model approaching to the upper bound of residual error. Besides, to enhance the performance of the FT-ZNN model, a fuzzy finite-time ZNN (FFT-ZNN), which possesses a fuzzy parameter, is further presented for solving the TVQPEI problem. A simulative example about the FT-ZNN and FFT-ZNN models solving the TVQPEI problem is given, and the experimental results expectably conform to the theoretical analysis. In addition, the designed FT-ZNN model is effectually applied to the repetitive motion of the three-link redundant robot and image fusion to show its potential practical value. Lin Xiao 0002, Lei Jia 0001, Yaonan Wang 0001, Jianhua Dai 0003, Qing Liao 0001, Quanxin Zhu |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2022 | Adams-Bashforth-Type Discrete-Time Zeroing Neural Networks Solving Time-Varying Complex Sylvester Equation With Enhanced RobustnessabstractIn this article, two Adams–Bashforth-type integration-enhanced discrete-time zeroing neural dynamic (ADTIZD) models are proposed to solve the time-varying complex Sylvester equation (TVCSE) problem in the first time. In ADTIZD models, Adams–Bashforth discrete formulas as novel discrete formulas are used, giving our ADTIZD models higher accuracy [truncation error being$O(\tau ^{5})$] but less time and space complexity than the ordinary multi-instant models. Enhanced by the integration part, the ADTIZD models can resist large additive noises, where even constant noises cannot decrease their precision. All convergence and robustness performance conclusions about our ADTIZD models are supported by rigorous theoretical proofs and numerical experiments. More comparisons between ADTIZD models and other discrete-time zeroing neural network models are shown in these experiments too. The efficacy of ADTIZD models is finally been validated in the simulation of adopting them in controlling a robotic manipulator. Zeshan Hu, Kenli Li 0001, Lin Xiao 0002, Yaonan Wang 0001, Mingxing Duan, Keqin Li 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 4 |
| 2022 | Distributed Population Dynamics for Searching Generalized Nash Equilibria of Population Games With Graphical Strategy InteractionsabstractEvolutionary games and population dynamics are finding increasing applications in design learning and control protocols for a variety of resource allocation problems. The implicit requirement for full communication has been the main limitation of the evolutionary game dynamic approach in engineering tasks with various information constraints. This article intends to build population games and dynamics with both static and dynamical graphical communication structures. To this end, we formulate a population game model with graphical strategy interactions and derive its corresponding population dynamics. In particular, we first introduce the concept of generalized Nash equilibria for population games with graphical strategy interactions, and establish the equivalence between the set of generalized Nash equilibria and the set of rest points of its distributed population dynamics. Furthermore, the conditions for convergence to generalized Nash equilibrium and particularly to Nash equilibrium are obtained for the distributed population dynamics with both static and dynamical graphical structures. These results provide a new approach to design distributed Nash equilibrium seeking algorithms for population games with both static and dynamical communication networks, and hence, expand the applicability of the population game dynamics in the design of learning and control protocols under distributed circumstances. Shaolin Tan, Yaonan Wang 0001, Athanasios V. Vasilakos |
IEEE Trans. Syst. Man Cybern. Syst. | 2 |
| 2022 | On Generalized Zeroing Neural Network Under Discrete and Distributed Time Delays and Its Application to Dynamic Lyapunov EquationabstractZeroing neural network (ZNN), an effective method for tracking solutions of dynamic equations, has been developed and improved by various strategies, typically the application of nonlinear activation functions (AFs) and varying parameters (VPs). Unlike VPs, AFs applied in ZNN models act directly on real-time error. The processing unit of v needs to obtain neural state in real time. In the implementation process, highly nonlinear AFs become an important cause of time delays, which eventually leads to instability and oscillation. However, most studies focus on exploring new theoretically valid AFs to improve performance of ZNNs, while ignoring the adverse effects of highly nonlinear AFs. The nonlinearity of AFs requires us fully consider time-delay tolerance of ZNNs using nonlinear AFs, so as to ensure that the model is not unstable even when disturbed by time delays. In this work, delay-perturbed generalized ZNN (DP-GZNN) is proposed to investigate time-delay tolerance of generalized ZNN (G-ZNN) in solving dynamic Lyapunov equation. Considering the nonlinearity of AFs, two delay terms are elegantly added to G-ZNN and DP-GZNN is then derived. After rigorous mathematical derivations, sufficient conditions in a linear matrix inequality (LMI) manner are presented for global convergence of DP-GZNN. Through rich numerical experiments, hyperparameters involved in the analysis process are discussed in detail. Comparative simulations are also conducted to compare the ability of different ZNN models to resist time delays. It is worth to mention that this is the first time to consider the ability of G-ZNN to resist discrete and distributed time delays. Qiuyue Zuo, Kenli Li 0001, Lin Xiao 0002, Yaonan Wang 0001, Keqin Li 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 4 |
| 2021 | Free-form Description Guided 3D Visual Graph Network for Object Grounding in Point Cloudabstract3D object grounding aims to locate the most relevant target object in a raw point cloud scene based on a freeform language description. Understanding complex and diverse descriptions, and lifting them directly to a point cloud is a new and challenging topic due to the irregular and sparse nature of point clouds. There are three main challenges in 3D object grounding: to find the main focus in the complex and diverse description; to understand the point cloud scene; and to locate the target object. In this paper, we address all three challenges. Firstly, we propose a language scene graph module to capture the rich structure and long-distance phrase correlations. Secondly, we introduce a multi-level 3D proposal relation graph module to extract the object-object and object-scene co-occurrence relationships, and strengthen the visual features of the initial proposals. Lastly, we develop a description guided 3D visual graph module to encode global contexts of phrases and proposals by a nodes matching strategy. Extensive experiments on challenging benchmark datasets (ScanRefer [3] and Nr3D [42]) show that our algorithm outperforms existing state-of-the-art. Our code is available at https://github.com/PNXD/FFL-3DOG. Mingtao Feng, Liang Zhang 0010, Guangming Zhu 0001, Hui Zhang 0023, Yaonan Wang 0001, Ajmal Mian |
ICCV | 8 |
| 2021 | Multi-Expert Adversarial Attack Detection in Person Re-identification Using Context InconsistencyabstractThe success of deep neural networks (DNNs) has promoted the widespread applications of person re-identification (ReID). However, ReID systems inherit the vulnerability of DNNs to malicious attacks of visually in-conspicuous adversarial perturbations. Detection of adversarial attacks is, therefore, a fundamental requirement for robust ReID systems. In this work, we propose a Multi-Expert Adversarial Attack Detection (MEAAD) approach to achieve this goal by checking context inconsistency, which is suitable for any DNN-based ReID systems. Specifically, three kinds of context inconsistencies caused by adversarial attacks are employed to learn a detector for distinguishing the perturbed examples, i.e., a) the embedding distances between a perturbed query person image and its top-K retrievals are generally larger than those between a benign query image and its top-K retrievals, b) the embedding distances among the top-K retrievals of a perturbed query image are larger than those of a benign query image, c) the top-K retrievals of a benign query image obtained with multiple expert ReID models tend to be consistent, which is not preserved when attacks are present. Extensive experiments on the Market1501 and DukeMTMC-ReID datasets show that, as the first adversarial attack detection approach for ReID, MEAAD effectively detects various adversarial attacks and achieves high ROC-AUC (over 97.5%). Shasha Li 0001, Min Liu 0008, Yaonan Wang 0001, Amit K. Roy-Chowdhury |
ICCV | 4 |
| 2021 | Automatic Leaf Diseases Detection System Based on Multi-stage Recognition
Songyun Deng, Lekai Cheng, Wei Sun 0028, Yaonan Wang 0001, Qiaokang Liang |
ICIG (1) | 5 |
| 2021 | Semi-supervised Cloud Edge Collaborative Power Transmission Line Insulator Anomaly Detection Framework
Yanqing Yang, Jianxu Mao, Hui Zhang 0023, Yurong Chen 0003, Hang Zhong, Yaonan Wang 0001 |
ICIG (1) | 7 |
| 2021 | Lane-free Autonomous Intersection Management: A Batch-processing Framework Integrating Reservation-based and Planning-based MethodsabstractAutonomous intersection management (AIM) refers to planning the trajectories for multiple connected and automated vehicles (CAVs) when they traverse an unsignalized intersection cooperatively. As an extension of the conventional AIM, lane-free AIM allows the CAVs to adjust their velocities and paths flexibly within the intersection. Nominally, one needs to formulate a centralized optimal control problem (OCP) to describe the concerned lane-free AIM scheme, but solving such an intractably scaled problem is challenging. This work proposes a batch-processing framework, which divides the traffic flow into batches. The cooperative trajectories within one batch are planned by numerically solving a small-scale OCP; all the batches are managed via a reservation-based method following the first-come-first-serve policy. The proposed batch-processing framework aims to run as fast as a reservation-based method at the macro level while taking care of the cooperative driving quality at the micro level. The proposed method is validated via simulation and preliminary experiments. Bai Li 0002, Youmin Zhang 0001, Tankut Acarman, Yakun Ouyang, Cagdas Yaman, Yaonan Wang 0001 |
ICRA | 6 |
| 2021 | FlowDriveNet: An End-to-End Network for Learning Driving Policies from Image Optical Flow and LiDAR Point FlowabstractLearning driving policies using an end-to-end network has been proved a promising solution for autonomous driving. Due to the lack of a benchmark driver behavior dataset that contains both the visual and the LiDAR data, existing works solely focus on learning driving from visual sensors. Besides, most works are limited to predict steering angle yet neglect the more challenging vehicle speed control problem. In this paper, we propose a novel end-to-end network, FlowDriveNet, which takes advantages of sequential visual data and LiDAR data jointly to predict steering angle and vehicle speed. The main challenges of this problem are how to efficiently extract driving-related information from images and point clouds, and how to fuse them effectively. To tackle these challenges, we propose a concept of point flow and declare that image optical flow and LiDAR point flow are significant motion cues for driving policy learning. Specifically, we first create an enhanced dataset that consists of images, point clouds and corresponding human driver behaviors. Then, in FlowDriveNet, a deep but efficient visual feature extraction module and a point feature extraction module are utilized to extract spatial features from optical flow and point flow, respectively. Additionally, a novel temporal fusion and prediction module is designed to fuse temporal information from the extracted spatial feature sequences and predict vehicle driving commands. Finally, a series of ablation experiments verify the importance of optical flow and point flow and comparison experiments show that our flow-based method outperforms the existing image-based approaches on the task of driving policy learning. Shuai Wang 0018, Jiahu Qin, Yaonan Wang 0001 |
ICRA | 4 |
| 2021 | Deep residual deconvolutional networks for defocus blur detectionabstractAbstract Accurate defocus blur detection has instigated wide research interest for the last few years. However, it is still a meaningful yet challenging machine vision task, and most methods rely on prior knowledge. Convolutional neural networks have proved the huge success for different tasks within the computer vision, and machine learning flew. A simple yet effective method of defocus blur detection was proposed in this paper, which by applying the deep residual convolutional encoder‐decoder network. The aims of DRDN is to automatically generate pixel‐level predictions for defocus blur images, and reconstruct output detection results of the same size as the input, which by performing several deconvolution operations at multiple scales through the transposed convolution, and skip connection. Afterwards, we used the slide window detection strategy and traversed the input image with a certain stride. Experiments on challenging benchmarks of defocus blur detection show that our algorithm achieved state‐of‐the‐art performance, and powerfully balanced the detection accuracy, and detection time. Kai Zeng 0010, Yaonan Wang 0001, Jianxu Mao, Xianen Zhou |
IET Image Process. | 2 |
| 2021 | Neural network-based adaptive hybrid impedance control for electrically driven flexible-joint robotic manipulators with input saturation
Shuai Ding 0007, Jinzhu Peng, Hui Zhang 0023, Yaonan Wang 0001 |
Neurocomputing | 4 |
| 2021 | EdgeGAN: One-way mapping generative adversarial network based on the edge information for unpaired training set
Qiaokang Liang, Youcheng Lei, Wei Sun 0028, Yaonan Wang 0001, Dan Zhang 0006 |
J. Vis. Commun. Image Represent. | 6 |
| 2021 | Sequence-tracker: Multiple object tracking with sequence features in severe occlusion scene
Qiaokang Liang, Wei Sun 0028, Yaonan Wang 0001, Dan Zhang 0006 |
J. Vis. Commun. Image Represent. | 5 |
| 2021 | DeepSeed Local Graph Matching for Densely Packed Cells TrackingabstractThe tracking of densely packed plant cells across microscopy image sequences is very challenging, because their appearance change greatly over time. A local graph matching algorithm was proposed to track such cells by exploiting the tight spatial topology of neighboring cells, and then an iterative searching strategy was used to grow the correspondence from a seed cell pair. Thus, the performance of the existing tracking approach heavily relies on the robustness of finding seed cell pair. However, the existing local graph matching algorithm cannot guarantee the correctness of the seed cell pair, especially in unregistered image sequences or image sequences with large time intervals. In this paper, we propose a DeepSeed local graph matching model to find seed cell pair robustly, by combining local graph matching and CNN-based similarity learning, which uses cells' spatial-temporal contextual information and cell pairs' similarity information. The CNN-based similarity learning is designed to learn cells' deep feature and measure cell pairs' similarity. Compared with the existing plant cell matching methods, the experimental results show that the DeepSeed local graph matching method can track most cells in unregistered image sequences. Moreover, the DeepSeed tracking algorithm can accurately track cells across image sequences with large time intervals. Min Liu 0008, Yalan Liu, Weili Qian, Yaonan Wang 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 4 |
| 2021 | Exploiting Global Camera Network Constraints for Unsupervised Video Person Re-IdentificationabstractMany unsupervised approaches have been proposed recently for the video-based re-identification problem since annotations of samples across cameras are time-consuming. However, higher-order relationships across the entire camera network are ignored by these methods, leading to contradictory outputs when matching results from different camera pairs are combined. In this paper, we address the problem of unsupervised video-based re-identification by proposing a consistent cross-view matching (CCM) framework, in which global camera network constraints are exploited to guarantee the matched pairs are with consistency. Specifically, we first propose to utilize the first neighbor of each sample to discover relations among samples and find the groups in each camera. Additionally, a cross-view matching strategy followed by global camera network constraints is proposed to explore the matching relationships across the entire camera network. Finally, we learn metric models for camera pairs progressively by alternatively mining consistent cross-view matching pairs and updating metric models using these obtained matches. Rigorous experiments on two widely-used benchmarks for video re-identification demonstrate the superiority of the proposed method over current state-of-the-art unsupervised methods; for example, on the MARS dataset, our method achieves an improvement of 4.2% over unsupervised methods, and even 2.5% over one-shot supervision-based methods for rank-1 accuracy. Rameswar Panda, Min Liu 0008, Yaonan Wang 0001, Amit K. Roy-Chowdhury |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2021 | Consensus With Persistently Exciting Couplings and Its Application to Vision-Based EstimationabstractThe problem of consensus in networked agent systems is revisited and applied to vision-based localization. A class of new consensus dynamics is introduced first, and sufficient conditions including the persistence of excitation on the coupling matrix for reaching consensus are derived. As an application of the proposed consensus dynamics, an adaptive localization algorithm then is proposed for autonomous robots equipped with primarily visual sensors in GPS-denied environments. In the context of consensus over an undirected tree topology, the convergence of the proposed localization algorithm is proved. Finally, both numerical simulations and physical experiments are presented to show the effectiveness of the proposed localization algorithm. Our algorithm is simpler to implement and computationally cheaper compared to other localization methods. Moreover, it is immune to error accumulation and long-term stable, and the asymptotical convergence of the estimation errors can be theoretically guaranteed. Zhiqiang Miao, Yun-Hui Liu 0001, Yaonan Wang 0001, Haoyao Chen, Hang Zhong, Rafael Fierro |
IEEE Trans. Cybern. | 3 |
| 2021 | Fixed-Time Adaptive Fuzzy Control for Uncertain Nonlinear SystemsabstractMost current methodologies on fixed-time adaptive fuzzy control for uncertain nonlinear systems result in practical fixed-time stability but not fixed-time stability, or require prior knowledge of the system dynamics. To obviate such restrictions, a fixed-time adaptive fuzzy control scheme is newly proposed with the discoveries of a singularity-avoidance virtual control design, a modified class of tuning functions and a projection operator-based adaptation mechanism. Fixed-time stability is established in the sense that the tracking error asymptotically converges to a user-defined interval within a prescribed fixed time. Illustrative examples verify the approaches developed. Kaixin Lu, Zhi Liu 0001, Yaonan Wang 0001, C. L. Philip Chen |
IEEE Trans. Fuzzy Syst. | 3 |
| 2021 | HDCB-Net: A Neural Network With the Hybrid Dilated Convolution for Pixel-Level Crack Detection on Concrete BridgesabstractCrack detection on concrete bridges is a critical task to ensure bridge safety. However, many cracks on concrete bridges show low contrast and blurry edges in practice, which brings challenges to image-based crack detection. In this article, to improve the detection accuracy of blurred cracks, we propose the HDCB-Net-a deep learning-based network with the hybrid dilated convolutional block (HDCB) for the pixel-level crack detection. Specifically, HDCB is employed to expand the receptive field of the convolution kernel without increasing the computational complexity and to avoid the gridding effect generated by the dilated convolution. Meanwhile, to achieve a reasonable efficiency/accuracy tradeoff, the HDCB-Net only contains a few downsampling stages, which can avoid the loss of blurred crack pixels due to excessive downsampling. Furthermore, a two-stage strategy is proposed to realize the fast crack detection in a massive number of images (more than 100 000) with the high resolution (5120 × 5120 pixels). At the first stage, YOLOv4 is employed to filter out images without cracks and generate coarse region proposals. At the second stage, to achieve refined damage analysis, the HDCB-Net is used to detect pixel-level cracks from the coarse region proposals. The experimental results demonstrate that the proposed HDCB-Net is genetic and able to improve the detection accuracy of blurred cracks, and our two-stage strategy is efficient for fast crack detection. The whole detection process takes only 0.64 s to handle a single image. Additionally, we have established a public dataset, including 150 632 high-resolution images, dedicated to the research of crack detection, which have been released along with this article. Wenbo Jiang 0002, Min Liu 0008, Yunuo Peng, Lehui Wu, Yaonan Wang 0001 |
IEEE Trans. Ind. Informatics | 5 |
| 2021 | A Parameter-Changing and Complex-Valued Zeroing Neural-Network for Finding Solution of Time-Varying Complex Linear Matrix Equations in Finite TimeabstractFor solving complex-valued linear matrix equations with time-varying coefficients (CV-LME-TVC) in the complex field, this article proposes a parameter-changing and complex-valued zeroing neural network (PC-CVZNN) model through integrating a new parameter-changing function. As compared to previous complex-valued zeroing neural networks (CVZNNs) with fixed parameters and existing parameter-changing functions, the PC-CVZNN model can achieve superior performance due to the accelerated role of the new parameter-changing function. In parts of theoretical analysis, we take advantage of Lyapunov methodology to prove that the proposed PC-CVZNN model can acquire the global and super-exponential convergence when the linear activation function is adopted, and even acquire super finite-time convergence when the new sign-bi-power activation function and its modified one are used. In parts of numerical comparison experiments, it is shown that the PC-CVZNN model possesses faster convergence rate than fixed-parameter CVZNN models and other analogy neural networks with parameter-changing function, when applied to finding the solution of CV-LME-TVC. Importantly, an application of the proposed method to the mobile manipulator control provides the potential practical value of the PC-CVZNN model in the industrial field. Lin Xiao 0002, Juan Tao, Jianhua Dai 0003, Yaonan Wang 0001, Lei Jia 0001, Yongjun He 0001 |
IEEE Trans. Ind. Informatics | 4 |
| 2021 | Relation Graph Network for 3D Object Detection in Point CloudsabstractConvolutional Neural Networks (CNNs) have emerged as a powerful tool for object detection in 2D images. However, their power has not been fully realised for detecting 3D objects directly in point clouds without conversion to regular grids. Moreover, existing state-of-the-art 3D object detection methods aim to recognize objects individually without exploiting their relationships during learning or inference. In this article, we first propose a strategy that associates the predictions of direction vectors with pseudo geometric centers, leading to a win-win solution for 3D bounding box candidates regression. Secondly, we propose point attention pooling to extract uniform appearance features for each 3D object proposal, benefiting from the learned direction features, semantic features and spatial coordinates of the object points. Finally, the appearance features are used together with the position features to build 3D object-object relationship graphs for all proposals to model their co-existence. We explore the effect of relation graphs on proposals' appearance feature enhancement under supervised and unsupervised settings. The proposed relation graph network comprises a 3D object proposal generation module and a 3D relation module, making it an end-to-end trainable network for detecting 3D objects in point clouds. Experiments on challenging benchmark point cloud datasets (SunRGB-D, ScanNet and KITTI) show that our algorithm performs better than existing state-of-the-art. Mingtao Feng, Syed Zulqarnain Gilani, Yaonan Wang 0001, Liang Zhang 0010, Ajmal Mian |
IEEE Trans. Image Process. | 3 |
| 2021 | AutoPedestrian: An Automatic Data Augmentation and Loss Function Search Scheme for Pedestrian DetectionabstractPedestrian detection is a challenging and hot research topic in the field of computer vision, especially for the crowded scenes where occlusion happens frequently. In this paper, we propose a novel AutoPedestrian scheme that automatically augments the pedestrian data and searches for suitable loss functions, aiming for better performance of pedestrian detection especially in crowded scenes. To our best knowledge, it is the first work to automatically search the optimal policy of data augmentation and loss function jointly for the pedestrian detection. To achieve the goal of searching the optimal augmentation scheme and loss function jointly, we first formulate the data augmentation policy and loss function as probability distributions based on different hyper-parameters. Then, we apply a double-loop scheme with importance-sampling to solve the optimization problem of data augmentation and loss function types efficiently. Comprehensive experiments on two popular benchmarks of CrowdHuman and CityPersons show the effectiveness of our proposed method. In particular, we achieve 40.58% in MR on CrowdHuman datasets and 11.3% in MR on CityPersons reasonable subset, yielding new state-of-the-art results on these two datasets. Baopu Li, Min Liu 0008, Yaonan Wang 0001, Wanli Ouyang |
IEEE Trans. Image Process. | 5 |
| 2021 | Learning Person Re-Identification Models From Videos With Weak SupervisionabstractMost person re-identification methods, being supervised techniques, suffer from the burden of massive annotation requirement. Unsupervised methods overcome this need for labeled data, but perform poorly compared to the supervised alternatives. In order to cope with this issue, we introduce the problem of learning person re-identification models from videos with weak supervision. The weak nature of the supervision arises from the requirement of video-level labels, i.e. person identities who appear in the video, in contrast to the more precise frame-level annotations. Towards this goal, we propose a multiple instance attention learning framework for person re-identification using such video-level labels. Specifically, we first cast the video person re-identification task into a multiple instance learning setting, in which person images in a video are collected into a bag. The relations between videos with similar labels can be utilized to identify persons, on top of that, we introduce a co-person attention mechanism which mines the similarity correlations between videos with person identities in common. The attention weights are obtained based on all person images instead of person tracklets in a video, making our learned model less affected by noisy annotations. Extensive experiments demonstrate the superiority of the proposed method over the related methods on two weakly labeled person re-identification datasets. Min Liu 0008, Dripta S. Raychaudhuri, Sujoy Paul, Yaonan Wang 0001, Amit K. Roy-Chowdhury |
IEEE Trans. Image Process. | 5 |
| 2021 | Efficient 3D Junction Detection in Biomedical Images Based on a Circular Sampling Model and Reverse MappingabstractDetection and localization of terminations and junctions is a key step in the morphological reconstruction of tree-like structures in images. Previously, a ray-shooting model was proposed to detect termination points automatically. In this paper, we propose an automatic method for 3D junction points detection in biomedical images, relying on a circular sampling model and a 2D-to-3D reverse mapping approach. First, the existing ray-shooting model is improved to a circular sampling model to extract the pixel intensity distribution feature across the potential branches around the point of interest. The computation cost can be reduced dramatically compared to the existing ray-shooting model. Then, the Density-Based Spatial Clustering of Applications with Noise (DBSCAN) algorithm is employed to detect 2D junction points in maximum intensity projections (MIPs) of sub-volume images in a given 3D image, by determining the number of branches in the candidate junction region. Further, a 2D-to-3D reverse mapping approach is used to map these detected 2D junction points in MIPs to the 3D junction points in the original 3D images. The proposed 3D junction point detection method is implemented as a build-in tool in the Vaa3D platform. Experiments on multiple 2D images and 3D images show average precision and recall rates of 87.11% and 88.33% respectively. In addition, the proposed algorithm is dozens of times faster than the existing deep-learning based model. The proposed method has excellent performance in both detection precision and computation efficiency for junction detection even in large-scale biomedical images. Lan Shen, Min Liu 0008, Chao Wang 0072, Changhao Guo, Erik Meijering, Yaonan Wang 0001 |
IEEE J. Biomed. Health Informatics | 6 |
| 2021 | Neuron Image Segmentation via Learning Deep Features and Enhancing Weak Neuronal StructuresabstractNeuron morphology reconstruction (tracing) in 3D volumetric images is critical for neuronal research. However, most existing neuron tracing methods are not applicable in challenging datasets where the neuron images are contaminated by noises or containing weak filament signals. In this paper, we present a two-stage 3D neuron segmentation approach via learning deep features and enhancing weak neuronal structures, to reduce the impact of image noise in the data and enhance the weak-signal neuronal structures. In the first stage, we train a voxel-wise multi-level fully convolutional network (FCN), which specializes in learning deep features, to obtain the 3D neuron image segmentation maps in an end-to-end manner. In the second stage, a ray-shooting model is employed to detect the discontinued segments in segmentation results of the first-stage, and the local neuron diameter of the broken point is estimated and direction of the filamentary fragment is detected by rayburst sampling algorithm. Then, a Hessian-repair model is built to repair the broken structures, by enhancing weak neuronal structures in a fibrous structure determined by the estimated local neuron diameter and the filamentary fragment direction. Experimental results demonstrate that our proposed segmentation approach achieves better segmentation performance than other state-of-the-art methods for 3D neuron segmentation. Compared with the neuron reconstruction results on the segmented images produced by other segmentation methods, the proposed approach gains 47.83% and 34.83% improvement in the average distance scores. The average Precision and Recall rates of the branch point detection with our proposed method are 38.74% and 22.53% higher than the detection results without segmentation. Bo Yang 0065, Weixun Chen, Huiqiong Luo, Yinghui Tan, Min Liu 0008, Yaonan Wang 0001 |
IEEE J. Biomed. Health Informatics | 6 |
| 2021 | Spherical-Patches Extraction for Deep-Learning-Based Critical Points Detection in 3D Neuron Microscopy ImagesabstractDigital reconstruction of neuronal structures is very important to neuroscience research. Many existing reconstruction algorithms require a set of good seed points. 3D neuron critical points, including terminations, branch points and cross-over points, are good candidates for such seed points. However, a method that can simultaneously detect all types of critical points has barely been explored. In this work, we present a method to simultaneously detect all 3 types of 3D critical points in neuron microscopy images, based on a spherical-patches extraction (SPE) method and a 2D multi-stream convolutional neural network (CNN). SPE uses a set of concentric spherical surfaces centered at a given critical point candidate to extract intensity distribution features around the point. Then, a group of 2D spherical patches is generated by projecting the surfaces into 2D rectangular image patches according to the orders of the azimuth and the polar angles. Finally, a 2D multi-stream CNN, in which each stream receives one spherical patch as input, is designed to learn the intensity distribution features from those spherical patches and classify the given critical point candidate into one of four classes: termination, branch point, cross-over point or non-critical point. Experimental results confirm that the proposed method outperforms other state-of-the-art critical points detection methods. The critical points based neuron reconstruction results demonstrate the potential of the detected neuron critical points to be good seed points for neuron reconstruction. Additionally, we have established a public dataset dedicated for neuron critical points detection, which has been released along with this article. Weixun Chen, Min Liu 0008, Qi Zhan, Yinghui Tan, Erik Meijering, Miroslav Radojevic, Yaonan Wang 0001 |
IEEE Trans. Medical Imaging | 7 |
| 2021 | MAMA Net: Multi-Scale Attention Memory Autoencoder Network for Anomaly DetectionabstractAnomaly detection refers to the identification of cases that do not conform to the expected pattern, which takes a key role in diverse research areas and application domains. Most of existing methods can be summarized as anomaly object detection-based and reconstruction error-based techniques. However, due to the bottleneck of defining encompasses of real-world high-diversity outliers and inaccessible inference process, individually, most of them have not derived groundbreaking progress. To deal with those imperfectness, and motivated by memory-based decision-making and visual attention mechanism as a filter to select environmental information in human vision perceptual system, in this paper, we propose a Multi-scale Attention Memory with hash addressing Autoencoder network (MAMA Net) for anomaly detection. First, to overcome a battery of problems result from the restricted stationary receptive field of convolution operator, we coin the multi-scale global spatial attention block which can be straightforwardly plugged into any networks as sampling, upsampling and downsampling function. On account of its efficient features representation ability, networks can achieve competitive results with only several level blocks. Second, it's observed that traditional autoencoder can only learn an ambiguous model that also reconstructs anomalies "well" due to lack of constraints in training and inference process. To mitigate this challenge, we design a hash addressing memory module that proves abnormalities to produce higher reconstruction error for classification. In addition, we couple the mean square error (MSE) with Wasserstein loss to improve the encoding data distribution. Experiments on various datasets, including two different COVID-19 datasets and one brain MRI (RIDER) dataset prove the robustness and excellent generalization of the proposed MAMA Net. Yurong Chen 0003, Hui Zhang 0023, Yaonan Wang 0001, Yimin Yang 0001, Xianen Zhou, Q. M. Jonathan Wu |
IEEE Trans. Medical Imaging | 3 |
| 2021 | 3D Neuron Microscopy Image Segmentation via the Ray-Shooting Model and a DC-BLSTM NetworkabstractThe morphology reconstruction (tracing) of neurons in 3D microscopy images is important to neuroscience research. However, this task remains very challenging because of the low signal-to-noise ratio (SNR) and the discontinued segments of neurite patterns in the images. In this paper, we present a neuronal structure segmentation method based on the ray-shooting model and the Long Short-Term Memory (LSTM)-based network to enhance the weak-signal neuronal structures and remove background noise in 3D neuron microscopy images. Specifically, the ray-shooting model is used to extract the intensity distribution features within a local region of the image. And we design a neural network based on the dual channel bidirectional LSTM (DC-BLSTM) to detect the foreground voxels according to the voxel-intensity features and boundary-response features extracted by multiple ray-shooting models that are generated in the whole image. This way, we transform the 3D image segmentation task into multiple 1D ray/sequence segmentation tasks, which makes it much easier to label the training samples than many existing Convolutional Neural Network (CNN) based 3D neuron image segmentation methods. In the experiments, we evaluate the performance of our method on the challenging 3D neuron images from two datasets, the BigNeuron dataset and the Whole Mouse Brain Sub-image (WMBS) dataset. Compared with the neuron tracing results on the segmented images produced by other state-of-the-art neuron segmentation methods, our method improves the distance scores by about 32% and 27% in the BigNeuron dataset, and about 38% and 27% in the WMBS dataset. Weixun Chen, Min Liu 0008, Yaonan Wang 0001, Erik Meijering |
IEEE Trans. Medical Imaging | 4 |
| 2021 | Discriminative Face Hallucination via Locality-Constrained and Category Embedding RepresentationabstractRecent years have witnessed the rapid development of face image hallucination techniques. However, the previous face hallucination methods are unsupervised and ignore the label information of training samples, leading to undesirable results. This article proposes a locality-constrained and category embedding representation (LCER) method to super-resolve face image in a supervised manner by embedding the label information in data representation. The proposed LCER incorporates the locality prior and category information into one unified framework, which aims to learn both the advantages of locality in preserving the true typologic structure of data manifold and the discriminability in exposing the class subspace information. Such strategy allows the LCER not only to preserve more sharpen image details but also to guarantee the face structure pattern be transferred mainly from the same subject in super-resolution reconstruction. Extensive experiments were conducted to evaluate the proposed LCER, and the comparative results demonstrate that it achieved superior face hallucination performance in both the quantitative measurements and visual impressions compared to several state-of-the-art. Licheng Liu, Rushi Lan, Yaonan Wang 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 3 |
| 2020 | Design and Analysis of a Synergy-Inspired Three-Fingered HandabstractHand synergy from neuroscience provides an effective tool for anthropomorphic hands to realize versatile grasping with simple planning and control. This paper aims to extend the synergy-inspired design from anthropomorphic hands to multi-fingered robot hands. The synergy-inspired hands are not necessarily humanoid in morphology but perform primary characteristics and functions similar to the human hand. At first, the biomechanics of hand synergy is investigated. Three biomechanical characteristics of the human hand synergy are explored as a basis for the mechanical simplification of the robot hands. Secondly, according to the synergy characteristics, a three-fingered hand is designed, and its kinematic model is developed for the analysis of some typical grasping and manipulation functions. Finally, a prototype is developed and preliminary grasping experiments validate the effectiveness of the design and analysis. Wenrui Chen, Zhilan Xiao, Jingwen Lu, Yaonan Wang 0001 |
ICRA | 5 |
| 2020 | Integrating Deformable Convolution and Pyramid Network in Cascade R-CNN for Fabric Defect DetectionabstractDefects on the surface of fabrics seriously affect the production speed and quality of textile products. There are many difficulties in the detection of surface defects on fabrics, such as substantial differences in length-width ratio, uneven distribution, and few features. However, existing methods have the disadvantages of slow detection speed and high misdetection rate. This present study proposes a method of integrating deformable convolution and pyramid network in Cascade R-CNN (IDPNet) for fabric defect detection. First, image data are labeled according to the type and distribution of defects. Then we design a novel multi-stage object detection architecture named IDPNet to detect defects on the surface of fabrics. In the first stage, Resnet50, in combination with feature pyramid network and deformable convolution is used to improve the detection performance of small defects. Besides, we trained a sequence of detectors with increasing IoUs stage by stage based on Cascade R-CNN in the second stage. Finally, experimental results demonstrate that the proposed neural network equip an outstanding performance against other approaches and achieve the accuracy of 91.57% in fabric defect detection, which proves its utility in practice. Honghao Li, Hui Zhang 0023, Li Liu 0060, Hang Zhong, Yaonan Wang 0001, Q. M. Jonathan Wu |
SMC | 5 |
| 2020 | Adaptive gradient-based block compressive sensing with sparsity for noisy images
Paul L. Rosin, Yukun Lai, Jinhua Zheng, Yaonan Wang 0001 |
Multim. Tools Appl. | 5 |
| 2020 | Autonomous mobile robot path planning in unknown dynamic environments using neural dynamics
Jiacheng Liang, Yaonan Wang 0001, Qi Pan, Jianhao Tan, Jianxu Mao |
Soft Comput. | 3 |
| 2020 | A Surface Defect Detection Framework for Glass Bottle Bottom Using Visual Attention Model and Wavelet TransformabstractGlass bottles must be thoroughly inspected before they are used for packaging. However, the vision inspection of bottle bottoms for defects remains a challenging task in quality control due to inaccurate localization, the difficulty in detecting defects in the texture region, and the intrinsically nonuniform brightness across the central panel. To overcome these problems, we propose a surface defect detection framework, which is composed of three main parts. First, a new localization method named entropy rate superpixel circle detection (ERSCD), which combines least-squares circle detection and entropy rate superpixel (ERS) with an improved randomized circle detection, is proposed to accurately obtain the region of interest (ROI) of the bottle bottom. Then, according to the structure-property, the ROI is divided into two measurement regions: central panel region and annular texture region. For the former, a defect detection method named frequency-tuned anisotropic diffusion super-pixel segmentation (FTADSP) that integrates frequency-tuned salient region detection (FT), anisotropic diffusion, and an improved superpixel segmentation is proposed to precisely detect the regions and boundaries of defects. For the latter, a defect detection strategy called wavelet transform multiscale filtering (WTMF) based on a wavelet transform and a multiscale filtering algorithm is proposed to reduce the influence of texture and to improve the robustness to localization error. The proposed framework is tested on four data sets obtained by our designed vision system. The experimental results demonstrate that our framework achieves the best performance compared with many traditional methods. Xianen Zhou, Yaonan Wang 0001, Qing Zhu 0003, Jianxu Mao, Changyan Xiao, Xiao Lu 0002, Hui Zhang 0023 |
IEEE Trans. Ind. Informatics | 2 |
| 2020 | DeepBranch: Deep Neural Networks for Branch Point Detection in Biomedical ImagesabstractMorphology reconstruction of tree-like structures in volumetric images, such as neurons, retinal blood vessels, and bronchi, is of fundamental interest for biomedical research. 3D branch points play an important role in many reconstruction applications, especially for graph-based or seed-based reconstruction methods and can help to visualize the morphology structures. There are a few hand-crafted models proposed to detect the branch points. However, they are highly dependent on the empirical setting of the parameters for different images. In this paper, we propose a DeepBranch model for branch point detection with two-level designed convolutional networks, a candidate region segmenter and a false positive reducer. On the first level, an improved 3D U-Net model with anisotropic convolution kernels is employed to detect initial candidates. Compared with the traditional sliding window strategy, the improved 3D U-Net can avoid massive redundant computations and dramatically speed up the detection process by employing dense-inference with fully convolutional neural networks (FCN). On the second level, a method based on multi-scale multi-view convolutional neural networks (MSMV-Net) is proposed for false positive reduction by feeding multi-scale views of 3D volumes into multiple streams of 2D convolution neural networks (CNNs), which can take full advantage of spatial contextual information as well as fit different sizes. Experiments on multiple 3D biomedical images of neurons, retinal blood vessels and bronchi confirm that the proposed 3D branch point detection method outperforms other state-of-the-art detection methods, and is helpful for graph-based or seed-based reconstruction methods. Yinghui Tan, Min Liu 0008, Weixun Chen, Hanchuan Peng, Yaonan Wang 0001 |
IEEE Trans. Medical Imaging | 6 |
| 2020 | Complex-Valued Discrete-Time Neural Dynamics for Perturbed Time-Dependent Complex Quadratic Programming With ApplicationsabstractIt has been reported that some specially designed recurrent neural networks and their related neural dynamics are efficient for solving quadratic programming (QP) problems in the real domain. A complex-valued QP problem is generated if its variable vector is composed of the magnitude and phase information, which is often depicted in a time-dependent form. Given the important role that complex-valued problems play in cybernetics and engineering, computational models with high accuracy and strong robustness are urgently needed, especially for time-dependent problems. However, the research on the online solution of time-dependent complex-valued problems has been much less investigated compared to time-dependent real-valued problems. In this article, to solve the online time-dependent complex-valued QP problems subject to linear constraints, two new discrete-time neural dynamics models, which can achieve global convergence performance in the presence of perturbations with the provided theoretical analyses, are proposed and investigated. In addition, the second proposed model is developed to eliminate the operation of explicit matrix inversion by introducing the quasi-Newton Broyden-Fletcher-Goldfarb-Shanno (BFGS) method. Moreover, computer simulation results and applications in robotics and filters are provided to illustrate the feasibility and superiority of the proposed models in comparison with the existing solutions. Yimeng Qi, Long Jin 0001, Yaonan Wang 0001, Lin Xiao 0002, Jiliang Zhang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2020 | Graphical Nash Equilibria and Replicator Dynamics on Complex NetworksabstractPairwise-interaction graphical games have been widely used in the study and design of strategic interaction in multiagent systems. With regard to this issue, one entitative problem is actually to understand how the interaction structure of agents affects the strategy configuration of Nash equilibria. This paper intends to study the effect of interaction networks on Nash equilibria in pairwise-interaction graphical games. We first show that interaction networks may induce new strategy equilibria in pairwise-interaction graphical games and then provide graphical conditions for the existence of these network-induced equilibria. Furthermore, to determine Nash equilibria of pairwise-interaction graphical games, a graphical replicator dynamics model is formulated, and its connection with graphical games is established. In detail, it is shown that every Nash equilibrium of the graphical games corresponds to a fixed point of the graphical replicator dynamics and that every asymptotically stable fixed point of the graphical replicator dynamics corresponds to a strict pure Nash equilibrium of the graphical games. The obtained results are applied in understanding coordination in complex networks and determination of structural conflicts in signed graphs. This work may provide new insights into understanding and designing strategy equilibria and dynamics in games on networks. Shaolin Tan, Yaonan Wang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2020 | Analysis and Synthesis of Underactuated Compliant Mechanisms Based on Transmission Properties of Motion and ForceabstractThis article analyzes and designs the transmission structure for underactuated compliant mechanisms (UCMs). The transmission structure of UCMs consists of serial and parallel transmission chains. At first, the UCMs are classified systematically according to the number and distribution of the serial and parallel transmissions. Next, the active and passive transmission properties of motion and force in UCMs are analyzed on the defined four subspaces of tangent and cotangent spaces of joint space. Synthesizing the classification and the transmission properties of UCMs, the congruent relationship between mechanical structure and transmission function is established, and different cases of UCMs are discussed and compared. A novel type of UCMs can achieve the independent regulation of passive stiffness, active force, and active motion that is useful for improving the transmission performance in robotic and prosthetic hands. Finally, a functional oriented design method is proposed and used to design a single-actuator two-fingered gripper for enveloping and precision grasps. The results demonstrate the validity of the proposed method. Wenrui Chen, Yaonan Wang 0001 |
IEEE Trans. Robotics | 3 |
| 2020 | Automatic semantic style transfer using deep convolutional neural networks and soft masks
Paul L. Rosin, Yukun Lai, Yaonan Wang 0001 |
Vis. Comput. | 4 |
| 2019 | A neural-network enhanced modeling method for real-time evaluation of the temperature distribution in a data center
Qiu Fang, Zhe Li 0050, Yaonan Wang 0001, Mengxuan Song, Jun Wang 0025 |
Neural Comput. Appl. | 3 |
| 2019 | Recurrent fuzzy wavelet neural networks based on robust adaptive sliding mode control for industrial robot manipulators
Vu Thi Yen, Yaonan Wang 0001, Cuong Van Pham |
Neural Comput. Appl. | 2 |
| 2019 | A Local Metric for Defocus Blur Detection Based on CNN Feature LearningabstractDefocus blur detection is an important and challenging task in computer vision and digital imaging fields. Previous work on defocus blur detection has put a lot of effort into designing local sharpness metric maps. This paper presents a simple yet effective method to automatically obtain the local metric map for defocus blur detection, which based on the feature learning of multiple convolutional neural networks (ConvNets). The ConvNets automatically learn the most locally relevant features at the super-pixel level of the image in a supervised manner. By extracting convolution kernels from the trained neural network structures and processing it with principal component analysis, we can automatically obtain the local sharpness metric by reshaping the principal component vector. Meanwhile, an effective iterative updating mechanism is proposed to refine the defocus blur detection result from coarse to fine by exploiting the intrinsic peculiarity of the hyperbolic tangent function. The experimental results demonstrate that our proposed method consistently performed better than the previous state-of-the-art methods. Kai Zeng 0010, Yaonan Wang 0001, Jianxu Mao, Junyang Liu, Weixing Peng, Nankai Chen |
IEEE Trans. Image Process. | 2 |
| 2019 | Weakly Supervised Biomedical Image Segmentation by Reiterative LearningabstractRecent advances in deep learning have produced encouraging results for biomedical image segmentation; however, outcomes rely heavily on comprehensive annotation. In this paper, we propose a neural network architecture and a new algorithm, known as overlapped region forecast, for the automatic segmentation of gastric cancer images. To the best of our knowledge, this report for the first time describes that deep learning has been applied to the segmentation of gastric cancer images. Moreover, a reiterative learning framework that achieves superior performance without pretraining or further manual annotation is presented to train a simple network on weakly annotated biomedical images. We customize the loss function to make the model converge faster while avoiding becoming trapped in local minima. Patch boundary errors were eliminated by our overlapped region forecast algorithm. By studying the characteristics of the model trained using two different patch extraction methods, we train iteratively and integrate predictions and weak annotations to improve the quality of the training data. Using these methods, a mean Intersection over Union coefficient of 0.883 and a mean accuracy of 91.09% were achieved on the partially labeled dataset, thereby securing a win in the 2017 China Big Data and Artificial Intelligence Innovation and Entrepreneurship Competition. Qiaokang Liang, Yang Nan 0002, Gianmarc Coppola, Kunglin Zou, Wei Sun 0028, Dan Zhang 0006, Yaonan Wang 0001, Guanzhen Yu |
IEEE J. Biomed. Health Informatics | 7 |
| 2019 | SSG: superpixel segmentation and GrabCut-based salient object segmentation
Xianen Zhou, Yaonan Wang 0001, Qing Zhu 0003, Changyan Xiao, Xiao Lu 0002 |
Vis. Comput. | 2 |
| 2018 | 3D Face Reconstruction from Light Field Images: A Model-Free Approach
Mingtao Feng, Syed Zulqarnain Gilani, Yaonan Wang 0001, Ajmal Mian |
ECCV (10) | 3 |
| 2018 | Methods and datasets on semantic segmentation: A review
Hongshan Yu, Zhengeng Yang, Yaonan Wang 0001, Wei Sun 0028, Mingui Sun, Yandong Tang |
Neurocomputing | 4 |
| 2018 | Distributed Estimation and Control for Leader-Following Formations of Nonholonomic Mobile RobotsabstractThe problem of the leader-following formation control of nonholonomic mobile robots is addressed in this paper. A distributed formation control strategy using explicitly the coordination errors among robots is proposed without assuming that each follower robot knows the full state of the leader. First, a distributed estimation law is proposed for each follower robot to estimate the states, including the position, orientation and linear velocity of the leader. The distributed formation control law is then designed based on the estimated states of the leader and the neighborhood formation tracking error. Under some mild assumptions on the interaction graph among the leader and the follower robots and the velocity of the leader, asymptotic convergence of formation tracking errors to zero can be achieved. Finally, some numerical simulations and experiments on a group of nonholonomic mobile robots are presented to demonstrate the effectiveness of the proposed strategy. Zhiqiang Miao, Yun-Hui Liu 0001, Yaonan Wang 0001, Rafael Fierro |
IEEE Trans Autom. Sci. Eng. | 3 |
| 2018 | Distributed Coordination of Islanded Microgrid Clusters Using a Two-Layer Intermittent Communication NetworkabstractThis paper proposes a distributed hierarchical cooperative control strategy for a cluster of islanded microgrids (MGs) with intermittent communication, which can regulate the frequency/voltage of all distributed generators (DGs) within each MG as well as ensure the active/reactive power sharing among MGs. A droop-based distributed secondary control scheme and a distributed tertiary control scheme are presented based on the iterative learning mechanics, by which the control inputs are merely updated at the end of each round of iteration, and thus, each DG only needs to share information with its neighbors intermittently in a low-bandwidth communication manner. A two-layer sparse communication network is modeled by pinning one or some DGs (pinned DGs) from the lower network of each MG to constitute an upper network. Under this control framework, the tertiary level generates the frequency/voltage references based on the active/reactive power mismatch among MGs while the pinned DGs propagate these references to their neighbors in the secondary level, and the frequency/voltage nominal set points for each DG in the primary level can be finally adjusted based on the frequency/voltage errors. Stability analysis of the two-layer control system is given, and sufficient conditions on the upper bound of the sampling period ratio of the tertiary layer to the secondary layer are also derived. The proposed controllers are distributed, and thus, allow different numbers of heterogeneous DGs in each MG. The effectiveness of the proposed control methodology is verified by the simulation of an ac MG cluster in Simulink/SimPower Systems. Xiaoqing Lu, Jingang Lai, Xinghuo Yu 0001, Yaonan Wang 0001, Josep M. Guerrero |
IEEE Trans. Ind. Informatics | 4 |
| 2018 | Benchmark Data Set and Method for Depth Estimation From Light Field ImagesabstractConvolutional Neural Networks (CNN) have performed extremely well for many image analysis tasks. However, supervised training of deep CNN architectures requires huge amounts of labelled data which is unavailable for light field images. In this paper, we leverage on synthetic light field images and propose a two stream CNN network that learns to estimate the disparities of multiple correlated neighbourhood pixels from their Epipolar Plane Images (EPI). Since the EPIs are unrelated except at their intersection, a two stream network is proposed to learn convolution weights individually for the EPIs and then combine the outputs of the two streams for disparity estimation. The CNN estimated disparity map is then refined using the central RGB light field image as a prior in a variational technique. We also propose a new real world dataset comprising light field images of 19 objects captured with the Lytro Illum camera in outdoor scenes and their corresponding 3D pointclouds, as ground truth, captured with the 3dMD scanner. This dataset will be made public to allow more precise 3D pointcloud level comparison of algorithms in the future which is currently not possible. Experiments on the synthetic and real world datasets show that our algorithm outperforms existing state-of-the-art for depth estimation from light field images. Mingtao Feng, Yaonan Wang 0001, Jian Liu 0014, Liang Zhang 0010, Hasan Firdaus M. Zaki, Ajmal Mian |
IEEE Trans. Image Process. | 2 |
| 2018 | Autoencoder With Invertible Functions for Dimension Reduction and Image ReconstructionabstractThe extreme learning machine (ELM), which was originally proposed for “generalized” single-hidden layer feedforward neural networks, provides efficient unified learning solutions for the applications of regression and classification. Although, it provides promising performance and robustness and has been used for various applications, the single-layer architecture possibly lacks the effectiveness when applied for natural signals. In order to over come this shortcoming, the following work indicates a new architecture based on multilayer network framework. The significant contribution of this paper are as follows: 1) unlike existing multilayer ELM, in which hidden nodes are obtained randomly, in this paper all hidden layers with invertible functions are calculated by pulling the network output back and putting it into hidden layers. Thus, the feature learning is enriched by additional information, which results in better performance; 2) in contrast to the existing multilayer network methods, which are usually efficient for classification applications, the proposed architecture is implemented for dimension reduction and image reconstruction; and 3) unlike other iterative learning-based deep networks (DL), the hidden layers of the proposed method are obtained via four steps. Therefore, it has much better learning efficiency than DL. Experimental results on 33 datasets indicate that, in comparison to the other existing dimension reduction techniques, the proposed method performs competitively better with fast training speeds. Yimin Yang 0001, Q. M. Jonathan Wu, Yaonan Wang 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 3 |
| 2017 | Matrix Separation Based on LMaFit-SeedabstractMatrix separation has a wide range of potential applications and many approaches have been devised to solve it. Especially, for applications involving large-scale data, such as vision tasks, improving the scalability of algorithms has attracted much attention. Reviewing these methods, they mainly involve convex optimization and factorization optimization. Convex optimization models become increasingly costly as the matrix size and rank grow. Factorization optimization models, to a large extent, reduce the computational complexity of matrix separation. l1-filtering applied the generalized Nyström to matrix separation, and proposed the seed-based convex optimization. In this paper, to benefit from the seed strategy, we propose matrix separation based on Low-Rank Matrix Fitting (LMaFit)-Seed, which is an algorithm on low-rank factorization optimization, to enhance the scalability in solving the problems of large-scale matrix separation and to be less time-consuming. We evaluate the proposed method and demonstrate comparisons with several state-of-the-art methods on synthetic data simulations and real sequences experiments. Hai-Xia Xu 0001, Wei Zhou 0027, Yaonan Wang 0001, Wei Wang 0025, Yan Mo |
Comput. J. | 3 |
| 2017 | Perception oriented transmission estimation for high quality image dehazing
Zhigang Ling, Guoliang Fan 0001, Jianwei Gong, Yaonan Wang 0001, Xiao Lu 0002 |
Neurocomputing | 4 |
| 2017 | A System for Automated Detection of Ampoule Injection ImpuritiesabstractAmpoule injection is a routinely used treatment in hospitals due to its rapid effect after intravenous injection. During manufacturing, tiny foreign particles can be present in the ampoule injection. Therefore, strict inspection must be performed before ampoule injections can be sold for hospital use. In the quality control inspection process, most ampoule enterprises still rely on manual inspection which suffers from inherent inconsistency and unreliability. This paper reports an automated system for inspecting foreign particles within ampoule injections. A custom-designed hardware platform is applied for ampoule transportation, particle agitation, and image capturing and analysis. Constructed trajectories of moving objects within liquid are proposed for use to differentiate foreign particles from air bubbles and random noise. To accurately classify foreign particles, multiple features including particle area, mean gray value, geometric invariant moments, and wavelet packet energy spectrum are used in supervised learning to generate feature vectors. The results show that the proposed algorithm is effective in classifying foreign particles and reducing false positive rates. The automated inspection system inspects over 150 ampoule injections per minute (versus ~ 12 ampoule injections per minute by technologist) with higher accuracy and repeatability. In addition, the automated system is capable of diagnosing impurity types while existing inspection systems are not able to classify detected particles. Ji Ge, Shaorong Xie, Yaonan Wang 0001, Jun Liu 0007, Hui Zhang 0023, Falu Weng, Changhai Ru, Chao Zhou 0002, Min Tan 0001, Yu Sun 0001 |
IEEE Trans Autom. Sci. Eng. | 3 |
| 2017 | Finite-Time Control for Robust Tracking Consensus in MASs With an Uncertain LeaderabstractThis paper investigates the finite-time control for robust tracking consensus problems of multiagent systems with an uncertain leader for situations where the state of the considered active leader may not be measured and the directed network topology is time-varying. Based on the neighbor-based state-estimation rule and a new Lyapunov stability analysis method, a continuous and nonlinear distributed tracking protocol using only relative position information is designed, under which each agent can follow the leader in finite time if the input (acceleration) of the leader is known, and the tracking errors can converge to a bounded region in finite time if the input of the leader is unknown. In particular, a special continuous distributed tracking protocol with bounded control inputs is introduced to track the active leader in finite time. Numerical simulations are also given to illustrate the effectiveness of the theoretic results. Xiaoqing Lu, Yaonan Wang 0001, Xinghuo Yu 0001, Jingang Lai |
IEEE Trans. Cybern. | 2 |
| 2017 | Evolutionary Dynamics of Collective Behavior Selection and Drift: Flocking, Collapse, and OscillationabstractBehavioral choice is ubiquitous across a wide range of interactive decision-making processes and a myriad of scientific disciplines. With regard to this issue, one entitative problem is actually to understand how collective social behaviors form and evolve among populations when they face a variety of conflict alternatives. In this paper, a selection-drift dynamic model is formulated to characterize the behavior imitation and exploration processes in social populations. Based on the proposed framework, several typical behavior evolution patterns, including behavioral flocking, collapse, and oscillation, are reproduced with different kinds of behavior networks. Interestingly, for the selection-drift dynamics on homogeneous symmetric behavior networks, we unveil the phase transition from behavioral flocking to collapse and derive the bifurcation diagram of the evolutionary stable behaviors in social behavior evolution. While via analyzing the survival conditions of the best behavior on heterogeneous symmetric behavior networks, we propose a selection-drift mechanism to guarantee consensus at the optimal behavior. Moreover, when the selection-drift dynamics on asymmetric behavior networks is simulated, it is shown that breaking the symmetry in behavior networks can induce various behavioral oscillations. These obtained results may shed new insights into understanding, detecting, and further controlling how social norm and cultural trends evolve. Shaolin Tan, Yaonan Wang 0001, Yao Chen 0003, Zhen Wang 0004 |
IEEE Trans. Cybern. | 2 |
| 2017 | Traffic Sign Recognition via Multi-Modal Tree-Structure Embedded Multi-Task LearningabstractTraffic sign recognition is a rather challenging task for intelligent transportation systems since signs in different subsets, e.g., speed limit signs, prohibition signs, and mandatory signs, are very different from each other in color or shape, whereas they share some similarities to the ones in the same subset. Therefore, it is important to integrate different modalities of visual features, such as color and shape, and select discriminative features for better sign description; in addition, it benefits to explore the correlations between the classes of traffic signs to learn the classifiers jointly to improve the generalization performance. In this paper, we propose Multi- Modal tree-structure embedded Multi-Task Learning called M2- tMTL to select discriminative visual features both between and within modalities, as well as the correlated features shared by similar classification tasks. Our method simultaneously introduces two structured sparsity-induced norms into a least squares regression. One of the norms can be used not only to select modality of features but also to conduct within-modality feature selection. Moreover, the hierarchical correlations among the classification tasks are well represented by a tree structure, and therefore, the tree-structure sparsity-induced norm is used for learning the regression coefficients jointly to boost the performance of multi-class traffic sign recognition. Alternating direction method of multipliers (ADMM) is used to efficiently solve the proposed model with guaranteed convergence. Extensive experiments on public benchmark data sets demonstrate that the proposed algorithm leads to a quite interpretable model, and it has better or competitive performance with several state-of-the-art methods but with less computational and memory cost. Xiao Lu 0002, Yaonan Wang 0001, Xuanyu Zhou, Zhenjun Zhang, Zhigang Ling |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2016 | Learning deep transmission network for single image dehazingabstractState-of-the-art single image dehazing algorithms have some challenges to deal with images captured under complex weather conditions because their assumptions usually do not hold in those situations. In this paper, we develop a deep transmission network for robust single image dehazing. This deep transmission network simultaneously copes with three color channels and local patch information to automatically explore and exploit haze-relevant features in a learning framework. We further explore different network structures and parameter settings to achieve tradeoffs between performance and speed, which shows that color channels information is the most useful haze-relevant feature rather than local information. Experiment results demonstrate that the proposed algorithm outperforms state-of-the-art methods on both synthetic and real-world datasets. Zhigang Ling, Guoliang Fan 0001, Yaonan Wang 0001, Xiao Lu 0002 |
ICIP | 3 |
| 2016 | Adaptive RBFNNs/integral sliding mode control for a quadrotor aircraft
Shushuai Li, Yaonan Wang 0001, Jianhao Tan |
Neurocomputing | 2 |
| 2016 | Adaptive trajectory tracking neural network control with robust compensator for robot manipulators
Cuong Van Pham, Yaonan Wang 0001 |
Neural Comput. Appl. | 2 |
| 2016 | A Method for Metric Learning with Multiple-Kernel Embedding
Xiao Lu 0002, Yaonan Wang 0001, Xuanyu Zhou, Zhigang Ling |
Neural Process. Lett. | 2 |
| 2016 | Erratum to: A Method for Metric Learning with Multiple-Kernel Embedding
Xiao Lu 0002, Yaonan Wang 0001, Xuanyu Zhou, Zhigang Ling |
Neural Process. Lett. | 2 |
| 2016 | Adaptive transmission compensation via human visual system for efficient single image dehazing
Zhigang Ling, Shutao Li 0001, Yaonan Wang 0001, He Shen 0001, Xiao Lu 0002 |
Vis. Comput. | 3 |
| 2015 | Minimum sweeping area motion planning for flexible serpentine surgical manipulator with kinematic constraintsabstractFlexible serpentine manipulators are widely used in surgical robots as it can be operated inside the patient's body cavity by backbone bending. However, during the bending the manipulator sweeps over a region, where sensitive organs may locate. This raises the safety concern. In this paper, a motion planning algorithm focusing on minimize the sweeping area for flexible serpentine manipulators is presented. Particularly, a three dimensional backward average neural dynamic model (BANDM) is proposed to build minimum sweeping area planning field in the configuration space of the serpentine manipulator. Given a target position, the motion sequence is generated automatically based on the established planning field. The simulations and experimental results validate the effectiveness and superiority of the proposed planning approach over conventional planning algorithms in terms of sweeping area with keeping target reach and obstacle avoidance. Zheng Li 0012, Wenjun Xu 0005, Yaonan Wang 0001, Hongliang Ren 0001 |
IROS | 4 |
| 2015 | Gradient-based compressive sensing for noise image and video reconstructionabstractIn this study, a fast gradient‐based compressive sensing (FGB‐CS) for noise image and video is proposed. Given a noise image or video, the authors first make it sparse by orthogonal transformation, and then reconstruct it by solving a convex optimisation problem with a novel gradient‐based method. The main contribution is twofold. Firstly, they deal with the noise signal reconstruction as a convex minimisation problem, and propose a new compressive sensing based on gradient‐based method for noise image and video. Secondly, to improve the computational efficiency of gradient‐based compressive sensing, they formulate the convex optimisation of noise signal reconstruction under Lipschitz gradient and replace the iteration parameter by the Lipschitz constant. With this strategy, the convergence of our FGB‐CS is reduced from O (1/ k ) to O (1/ k 2 ). Experimental results indicate that their FGB‐CS method is able to achieve better performance than several classical algorithms. Yaonan Wang 0001, Xiaojiang Peng |
IET Commun. | 2 |
| 2015 | Adaptive extended piecewise histogram equalisation for dark image enhancementabstractHistogram equalisation has been widely used for image enhancement because of its simple implementation and satisfactory performance. However, traditional histogram equalisation uniformly redistributes an entire histogram or multiple piecewise histograms with the same equalisation strategy, which may produce unnatural artefacts, over‐enhancement or under‐enhancement in wide dynamic range dark image enhancement. This study proposes an adaptive extended piecewise histogram equalisation algorithm (AEPHE) for dark image enhancement. First, an original histogram is divided into a group of extended piecewise histograms. Then, an adaptive histogram equalisation, which balances intensity preservation and contrast boosting, is further developed and respectively applied to these extended piecewise histograms. The final histogram for image enhancement is produced by a weighted fusion of these equalised histograms. The experimental results indicate that AEPHE is superior to multiple state‐of‐the‐art algorithms. Zhigang Ling, Yan Liang 0001, Yaonan Wang 0001, He Shen 0001, Xiao Lu 0002 |
IET Image Process. | 3 |
| 2015 | Data Partition Learning With Multiple Extreme Learning MachinesabstractAs demonstrated earlier, the learning accuracy of the single-layer-feedforward-network (SLFN) is generally far lower than expected, which has been a major bottleneck for many applications. In fact, for some large real problems, it is accepted that after tremendous learning time (within finite epochs), the network output error of SLFN will stop or reduce increasingly slowly. This report offers an extreme learning machine (ELM)-based learning method, referred to as the parent-offspring progressive learning method. The proposed method works by separating the data points into various parts, and then multiple ELMs learn and identify the clustered parts separately. The key advantages of the proposed algorithms as compared to the traditional supervised methods are twofold. First, it extends the ELM learning method from a single neural network to a multinetwork learning system, as the proposed multiELM method can approximate any target continuous function and classify disjointed regions. Second, the proposed method tends to deliver a similar or much better generalization performance than other learning methods. All the methods proposed in this paper are tested on both artificial and real datasets. Yimin Yang 0001, Q. M. Jonathan Wu, Yaonan Wang 0001, Zeeshan Khawar Malik, Xiaofang Yuan |
IEEE Trans. Cybern. | 3 |
| 2015 | A Method to Calibrate Vehicle-Mounted Cameras Under Urban Traffic ScenesabstractWe address the problem of vehicle-mounted camera calibration under urban traffic scenes regarding the fact that the traditional calibration methods are practically restricted, since the internal parameters should be calibrated in the laboratory and it is impossible for recalibration that resulted from the parameters drifting or re-focusing when driving on roads. In this paper, we propose to utilize the manual lines lying in Manhattan directions in the scenes to compute their corresponding vanishing points for camera calibration, as the urban traffic scenes are usually man-made and the important lines and signs for driving are typically lying in the Manhattan directions. For “Manhattan world” scenes, where there are plenty of lines lying in Manhattan directions, the lines in the scene are detected automatically, and the clusters corresponding to Manhattan directions are obtained using RANSAC-like methods. For the more general “quasi-Manhattan world” scenes, where only the lines in two directions can be found naturally, while the lines in the other direction are usually detected trivially or even can be hardly detected, we propose a method to estimate the lines in the third direction to improve the vanishing point estimation accuracy. The method proposed is tested on both two types of scenes, and the accuracy and practicability of this method are demonstrated. Furthermore, calibration experiments on both one image and multiple images are conducted, which show that the results can be more accurate when more images are used. Yaonan Wang 0001, Xiao Lu 0002, Zhigang Ling, Yimin Yang 0001, Zhenjun Zhang, Kena Wang |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2015 | Progressive Learning Machine: A New Approach for General Hybrid System ApproximationabstractAs the most important property of neural networks (NNs), the universal approximation capability of NNs is widely used in many applications. However, this property is generally proven for continuous systems. Most industrial systems are hybrid systems (e.g., piecewise continuous), which is a significant limitation for real applications. Recently, many identification methods have been proposed for hybrid system approximation; however, these methods only operate in linear hybrid systems. In this paper, the progressive learning machine-a new learning algorithm based on multi-NNs-is proposed for general hybrid nonlinear/linear system approximation. This algorithm classifies hybrid systems into several continuous systems and can approximate any hybrid system with zero output error. The performance of the proposed learning method is demonstrated via numerical examples and with experimental data from real applications. Yimin Yang 0001, Yaonan Wang 0001, Q. M. Jonathan Wu, Min Liu 0008 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |