EDBT 2026 Demo / reviewers in the wild / expert
Yue Wu 0004
dblp:41/5979-4
· DBLP profile ↗
97ranked-venue papers
17as first author
80since 2021 · last 2026
0000-0002-3459-5079ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 51 · 10 first-author · 44 since 2021Graphics, computer vision, multimedia, augmented reality and games · 29 · 7 first-author · 28 since 2021Applied, interdisciplinary, general and emerging computing · 21 · 2 first-author · 14 since 2021Databases, data management, data science and information retrieval · 5 · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Computer networks · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Hybrid Vector-Occupancy Field for Robust Implicit 3D Surface ReconstructionabstractWe introduce the Hybrid Vector-Occupancy Field (HVOF), a new implicit 3D representation for reconstructing both open and closed surfaces from sparse point clouds. Existing approaches, such as occupancy field and signed distance fields, face severe limitations. They struggle with open surfaces, while unsigned distance field and neural vector field exhibit directional instability in complex topologies and ridge regions. HVOF addresses these challenges by incorporating a smoothly decaying occupancy field around the surface, while capturing precise local geometry using truncated displacement vectors, naturally mitigating direction-field ambiguities near ridge regions. This unified design forms a robust hybrid representation that leverages both occupancy and vector fields. To fulfill it, we design a Hybrid Field variational autoencoder including a hierarchical cross-attention encoder and dual-branch decoder that jointly learn occupancy and vector fields through continuous weighting. Extensive experiments demonstrate that HVOF consistently outperforms state-of-the-art methods across ShapeNet, ABC, and MGN datasets, accurately reconstructing both open and closed surfaces while preserving fine geometric details in complex regions. Yue Wu 0004, Tengfei Xiao, Can Qin, Yongzhe Yuan, Hao Li 0009, Kaiyuan Feng, Wenping Ma 0001 |
AAAI | 1 |
| 2026 | DcSplat: Dual-Constraint Human Gaussian Splatting with Latent Multi-View ConsistencyabstractHuman Novel View Synthesis (HNVS) aims to synthesize photorealistic human images from novel viewpoints given observations from known views. Despite significant advances achieved by existing methods such as NeRF, diffusion models, and 3DGS, they still face substantial challenges in achieving stable modeling from a single image. In this paper, we introduce Dual-Constraint Human Gaussian Splatting (DcSplat), a novel, simple, and efficient 3D Gaussian-based framework for single-view 3D human reconstruction. To address occlusion-induced texture missing and depth ambiguities, we introduce two key components: a Latent Multi-View Consistency Constraint Mechanism and a Geometric Constraint Module. The former employs a Latent-space Appearance Transformer (LatentFormer) to learn semantically coherent, view-consistent appearance priors via SMPL-guided pseudo-view fusion. The latter refines noisy SMPL-based depth through a U-Net-like structure conditioned on latent appearance features. These two modules are jointly optimized to generate high-quality Gaussian parameters in a unified latent space. Extensive experiments demonstrate that DcSplat outperforms existing SOTA methods in both geometry and texture quality, while achieving fast inference and lower computational cost. Tengfei Xiao, Yue Wu 0004, Yongzhe Yuan, Can Qin, Hao Li 0009, Mingyang Zhang 0002 |
AAAI | 2 |
| 2026 | BCFNet: Bi-temporal collaborative fusion network for multi-modal humor detection
Boya Deng, Jianzhao Li, Maoguo Gong, Zedong Tang, Yourun Zhang, Kaiyuan Feng, Yue Wu 0004 |
Pattern Recognit. | 7 |
| 2026 | Sparse Unmixing Guided Adversarial Attack for Hyperspectral Image ClassificationabstractIn recent years, adversarial attacks in hyperspectral image (HSI) classification have garnered increasing attention. However, existing attack methods primarily manipulate individual pixel spectral to mislead deep neural networks (DNNs) into misclassification, overlooking the physical consistency of hyperspectral data. This oversight results in adversarial samples that lack physical interpretability and suffer from low attack efficiency. To alleviate these issues, this paper proposes a sparse unmixing guided adversarial attack framework (SUGAA) to efficiently generate hyperspectral adversarial samples that satisfy physical consistency. The proposed framework first employs sparse unmixing to extract the abundance matrix of HSI, introducing adversarial perturbations to the abundance matrix to generate physically consistent adversarial samples. Additionally, SUGAA leverages the compositional similarity of materials within intra-class HSI pixels to design a class-specific perturbation generation strategy, enhancing the applicability of adversarial perturbations across pixels of the same class. To further improve optimization effectiveness, SUGAA incorporates a class-specific perturbation optimization algorithm based on momentum iterative gradients to avoid local optima, ensuring stable and efficient perturbation generation. Experimental results on real HSI datasets demonstrate that SUGAA not only generates adversarial samples with high attack performance and physical consistency but also exhibits robustness to common preprocessing transformations. Hao Li 0009, Kelin Dang, Maoguo Gong, A. K. Qin 0001, Yu Zhou 0051, Yue Wu 0004, Lining Xing 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2026 | Multigranularity Adversarial Attacks on Large Language Models Using Genetic ProgrammingabstractLarge language models (LLMs) have demonstrated remarkable capabilities across various natural language processing tasks, but they remain vulnerable to adversarial attacks and pose significant security concerns. Existing attack methods often treat adversarial prompts as flat sequences, neglecting the rich hierarchical structure of natural language, which could limit their effectiveness. Advancing the methodologies for adversarial attacks is crucial for rigorously assessing the security of LLMs and identifying subtle vulnerabilities. This paper introduces AdvGP, a novel framework that leverages genetic programming (GP) to generate adversarial prompts for LLMs. AdvGP exploits the inherent structural similarities between GP trees and natural language syntax to optimize the structure of harmful prompts. The framework incorporates a multi-granularity hierarchical attack strategy, specialized genetic operators that leverage an assisting LLM for depth-aware crossover and multi-level mutation, and a comprehensive fitness function integrating semantic consistency and attack effectiveness. The proposed method achieves competitive attack performance on multiple LLMs, consistently generating harmful outputs despite higher perplexity than some baselines. Ablation studies confirm the significant contributions of both LLM-aided and depth-aware mechanisms to AdvGP’s effectiveness. Furthermore, transferability analysis reveals that the generated prompts are able to bypass the defenses of various state-of-the-art LLMs, such as ChatGPT and Gemini. Wencheng Han, Hao Li 0009, Maoguo Gong, Yu Zhou 0051, Yue Wu 0004, A. K. Qin 0001, Lining Xing 0001 |
IEEE Trans. Evol. Comput. | 5 |
| 2026 | Many-Problem Surrogates for Transfer Evolutionary Multiobjective Optimization With Sparse Transfer StackingabstractFor expensive multiobjective optimization problems, there exists useful knowledge, e.g., the trained surrogate models, can be transferred to assist the optimization of a target optimization problem, which is termed as multi-problem surrogates. Stacking transfer is able to combine the pretrained source surrogate models and the preliminary target model with a meta-regression algorithm to transfer knowledge from source to target. However, when large-scale source models are involved in the many-problem scenarios, the less correlated sources may hurt the target performance, which is known as negative transfer. In this paper, sparse representation of the coefficients of meta-regression is considered to automatically select the most relevant source models for largely avoiding negative transfer. In the proposed many-problem surrogates, the coefficients of the source and target models are assumed to be sparse under the non-negativity and sum-to-one constraints. Then a sparse transfer stacking model is established with l1-norm of the coefficients. Next, the alternating direction method of multipliers is employed to solve the resulting constrained optimization problem by converting it into several much simpler problems. Most of the previous works assume that the costs for evaluation have no much difference and this assumption rarely holds in the real-world applications. In order to further reduce the total costs, an improved surrogate model with a cost-sensitive measure is designed to estimate the cost and select new solutions for real evaluation based on their estimated fitness, uncertainty and cost. Experimental results on synthetic and practical problems have demonstrated the superiority of the proposed many-problem surrogates. Hao Li 0009, Fanggao Wan, Maoguo Gong, A. K. Qin 0001, Yue Wu 0004, Lining Xing 0001 |
IEEE Trans. Evol. Comput. | 5 |
| 2026 | Privacy-Enhanced Offline Data-Driven Evolutionary Optimization Based on Cloud ServerabstractData-driven evolutionary algorithms (DDEAs) have achieved significant success in numerous real-world optimization problems, where exact objective functions and constraint functions do not exist, and they mainly rely on available data. However, the existing DDEAs primarily focus on improving performance through data and surrogate, without considering that the users may lack the specialized domain knowledge and sufficient computing resources required for DDEAs. To address the aforementioned issues, this paper proposes a novel paradigm called Evolutionary Learning and Optimization as a Service (ELOaaS) and investigates the potential collusion attacks between machine learning modules and evolutionary computing modules on cloud server, which may lead to privacy leakage. Consequently, a privacy-enhanced DDEA (PEDDEA) is proposed as an instantiation algorithm of ELOaaS, which is designed to tackle offline data-driven evolutionary optimization within the ELOaaS paradigm. In the proposed PEDDEA, a subspace learning-based privacy protection strategy is designed to defense the collusion attacks. Additionally, a model management strategy based on Kendall tau metric is introduced to construct high-quality surrogate ensembles. PEDDEA enables users to outsource private offline data to cloud servers, thereby approaching the optimal solution while ensuring privacy protection. Comprehensive experiments are conducted on benchmark problems and safety evaluation problems of autonomous vehicles. According to the experimental results, the proposed algorithm has significant performance advantages over existing offline DDEAs while ensuring privacy protection. Hao Li 0009, Zhibin Xu, Maoguo Gong, A. K. Qin 0001, Yue Wu 0004, Lining Xing 0001, Yu Zhou 0051 |
IEEE Trans. Evol. Comput. | 5 |
| 2025 | MUCD: Unsupervised Point Cloud Change Detection via Masked Consistencyabstract3D Change Detection (3DCD) has gradually become another research hotspot after image change detection. Recent works focus on using artificial labels for supervised or weakly-supervised training of siamese networks to segment changed points. However, labeling every points of multi-temporal point clouds is very expensive and time-consuming. In addition, these works lack effective self-supervised signals, and existing self-supervised signals often fail to capture sufficiently rich change information. To solve this problem, we assume that the powerful representation of 3D objects should model the consistency information of unchanged regions and distinguish different objects. Based on this assumption, we propose a new unsupervised framework called MUCD to learn change information of multi-temporal point clouds through bidirectional optimization of change segmentor and feature extractor. The training of network is divided into two stages. We first design a foreknowledge point contrastive loss based on the characteristics of the 3DCD task to initialize the feature extractor, and then propose a masked consistency loss to further learn the shared geometric information of unchanged regions in the multi-temporal point clouds, utilizing it as a free and powerful supervised signal to train a change segmentor. In the inference stage, only the segmentor is used to take multi-temporal point clouds as input and produce change segmentation result. Extensive experiments are conducted on SLPCCD and Urb3DCD, two real-world datasets of streets and urban buildings, to verify that our proposed unsupervised method is highly competitive and even outperforms supervised methods in scenes where semantic information changes occur, exhibiting better performance in generalization ability and robustness. Yue Wu 0004, Yongzhe Yuan, Maoguo Gong, Hao Li 0009, Mingyang Zhang 0002, Wenping Ma 0001, Qiguang Miao |
AAAI | 1 |
| 2025 | AdvDisplay: Adversarial Display Assembled by Thermoelectric Cooler for Fooling Thermal Infrared DetectorsabstractWhen the current physical adversarial patches cannot deceive thermal infrared detectors, the existing techniques implement adversarial attacks from scratch, such as digital patch generation, material production, and physical deployment. Besides, it is difficult to finely regulate infrared radiation. To address these issues, this paper designs an adversarial thermal display (AdvDisplay ) by assembling thermoelectric coolers (TECs) as an array. Specifically, to reduce the gap between patches in the physical and digital worlds and decrease the power of AdvDisplay device, heat transfer loss and electric power loss are designed to guide the patch optimization. In addition, a precise temperature control scheme for AdvDisplay is proposed based on proportional-integral-derivative (PID) control. Due to the accurate temperature regulation and the reusability of AdvDisplay , our method is able to improve the attack success rate and the efficiency of physical deployments. Extensive experimental results indicate that the proposed method possesses superior adversarial effectiveness compared to other methods and demonstrates strong robustness in physical attacks. Hao Li 0009, Fanggao Wan, Yue Wu 0004, Mingyang Zhang 0002, Maoguo Gong |
AAAI | 4 |
| 2025 | Where Precision Meets Efficiency: Transformation Diffusion Model for Point Cloud RegistrationabstractWe propose a transformation diffusion model for point cloud registration to balance precision and efficiency. Our method formulates point cloud registration as a denoising diffusion process from noisy transformation to object transformation, which is represented by quaternion and translation. Specifically, in training stage, object transformation diffuses from ground-truth transformation to random distribution, and the model learns to reverse this noising process. In sampling stage, the model refines randomly generated transformation to the optimal transformation in a progressive way. We derive the variational bound in closed form for training and provide instantiation of the model. Our diffusion model maps transformation into latent space, and splits the transformation into two components (rotation and translation) based on the fact that they belong to different solution spaces. In addition, our work provides the following crucial findings: (i) Point cloud registration, one of the representative discriminative tasks, can be solved by a generative way and mapped into latent space to obtain new unified probabilistic formulation. (ii) Our model, Transformation Diffusion Model (TDM) can be a plug-and-play agent for point cloud registration, making our method applicable to different deep registration networks. Experimental results on synthetic and real-world datasets demonstrate that, in correspondence-free and correspondence-based scenarios, TDM can both achieve exceeding 60% performance improvements and higher efficiency simultaneously. Yongzhe Yuan, Yue Wu 0004, Xiaolong Fan, Maoguo Gong, Qiguang Miao, Wenping Ma 0001 |
AAAI | 2 |
| 2025 | Partial Point Cloud Registration with Multi-view 2D Image LearningabstractLearning representations from numerous 2D image data has shown promising performance, yet very few works apply this representations to point cloud registration. In this paper, we explore how to leverage the 2D information to assist the point cloud registration, and propose IAPReg, an Image-Assisted Partial 3D point cloud Registration framework with the multi-view images generated by the input point cloud. It is expected to enrich 3D information with 2D knowledge, and leverage 2D knowledge to assist with point cloud registration. Specifically, we create multi-view depth maps by projecting the input point cloud from several specific views, and then extract 2D and 3D features using some well-established models. To fuse the information learned from 2D and 3D modalities, inter-modality multi-view learning module is proposed to enhance geometric information and complement semantic information. Weighted SVD is a common method to reduce the impact of inaccurate correspondences on registration. However, determining the correspondence weights is not trivial. Therefore, we design a 2D-weighted SVD method, where the 2D knowledge is employed to provide weight information of correspondences. Extensive experiments perform that our method outperform the state-of-the-art method without additional 2D training data. Yue Zhang 0040, Yue Wu 0004, Wenping Ma 0001, Maoguo Gong, Hao Li 0009, Biao Hou |
AAAI | 2 |
| 2025 | Disentangled Pose and Appearance Guidance for Multi-Pose GenerationabstractHuman pose generation is a complex task due to the non-rigid and highly variable nature of human body structures and appearances. However, existing methods often overlook the fundamental differences between spatial transformations of poses and texture generation for appearance, which makes them prone to overfitting. To address this issue, we propose a multi-pose generation framework driven by disentangled pose and appearance guidance. Our approach includes a Global-aware Pose Generation module that iteratively generates pose embeddings, enabling effective control over non-rigid body deformations. Additionally, we introduce the Global-aware Transformer Decoder, which leverages similarity queries and attention mechanisms to achieve spatial transformations and enhance pose consistency through a Global-aware block. In the appearance generation phase, we condition a diffusion model on pose embeddings produced in the initial stage and introduce an Appearance Adapter that extracts high-level contextual semantic information from multi-scale features, enabling further refinement of pose appearance textures and providing appearance guidance. Extensive experiments on the UBC Fashion and TikTok datasets demonstrate that our framework achieves state-of-the-art results in both quality and fidelity, establishing it as a powerful approach for complex pose generation tasks. Tengfei Xiao, Yue Wu 0004, Can Qin, Maoguo Gong, Qiguang Miao, Wenping Ma 0001 |
CVPR | 2 |
| 2025 | Lite3D: A Lightweight Hybrid CNN-Transformer Framework for 3D Video StabilizationabstractVideo stabilization plays a critical role in IoT, as it enhances the accuracy and reliability of visual data collected from mobile and unstable devices, enabling more effective monitoring and analysis. However, existing stabilization methods frequently lack sufficient accuracy in motion estimation or struggle with expensive computation. In this paper, we introduce Lite3D, a novel lightweight 3D deep learning-based video stabilization framework that combines a hybrid CNN and Transformer architecture for more efficient feature extraction. During training, our method utilizes estimated depth maps and relative camera poses to generate target views, jointly optimizing DepthNet and PoseNet through an unsupervised learning strategy. Compared to existing 3D video stabilization methods, our approach enables more accurate learning of motion information. In the inference phase, we smooth the estimated camera pose trajectory and synthesize stabilized frames using the optimal depth maps. Experimental results on the NUS dataset demonstrate that Lite3D achieves state-of-the-art stability, with additional tests on the DeepStab dataset confirming strong generalization. Moreover, our method contains fewer parameters compared to existing 3D deep learning-based video stabilization, offering both high stabilization quality and computational efficiency. Zejing Shan, Yiping Duan, Yue Wu 0004, Qiyuan Du, Xiaoming Tao 0001 |
ICC | 3 |
| 2025 | PointTruss: K-Truss for Point Cloud RegistrationabstractPoint cloud registration is a fundamental task in 3D computer vision. Recent advances have shown that graph-based methods are effective for outlier rejection in this context. However, existing clique-based methods impose overly strict constraints and are NP-hard, making it difficult to achieve both robustness and efficiency. While the k-core reduces computational complexity, which only considers node degree and ignores higher-order topological structures such as triangles, limiting its effectiveness in complex scenarios. To overcome these limitations, we introduce the $k$-truss from graph theory into point cloud registration, leveraging triangle support as a constraint for inlier selection. We further propose a consensus voting-based low-scale sampling strategy to efficiently extract the structural skeleton of the point cloud prior to $k$-truss decomposition. Additionally, we design a spatial distribution score that balances coverage and uniformity of inliers, preventing selections that concentrate on sparse local clusters. Extensive experiments on KITTI, 3DMatch, and 3DLoMatch demonstrate that our method consistently outperforms both traditional and learning-based approaches in various indoor and outdoor scenarios, achieving state-of-the-art results. Yue Wu 0004, Yongzhe Yuan, Maoguo Gong, Qiguang Miao, Hao Li 0009, Mingyang Zhang 0002, Wenping Ma 0001 |
NeurIPS | 1 |
| 2025 | Multi-scale subspace co-clustering network with adaptive multi-scale enhancement for remote sensing scene classification
Zhenping Sun, Changyu Chen, Haiyan Han, Yue Wu 0004 |
Knowl. Based Syst. | 5 |
| 2025 | Revisiting Siamese-Based 3D Single Object Tracking With a Versatile Transformerabstract3D Single Object Tracking (SOT) plays an important role in real-world visual applications such as autonomous driving and planning. How to realize effective 3D SOT is still a valuable challenge due to its carrier-sparse point clouds and its role-complex influencing factors. Inspired by the remote modeling of popular transformers, we further propose a Versatile Point Tracking Transformer (VPTT) method for 3D SOT, with object guidance from the template point cloud to the search area point cloud under the siamese-based tracking paradigm. Specifically, VPTT employs self- and cross- attention mechanisms and extends four matching operations, resulting in leveraging the contextual information of consecutive frames to improve the tracking results. By constructing a deep network VerFormer consisting of four successive transformer layers, which performs matching operations involving fusional transformation, separative discrimination, intersectional interaction, and unidirectional propagation from shallow to deep. Considering that the tracking task involves multiple processes, VPTT further learns how to forecast intermediate outputs including mask probability, trailing distance, and heading angle at each stage. Such a specialized design allows our VPTT to revisit the end-to-end training paradigm used for 3D tracking while developing a versatile transformer that is a perfect fit for the 3D SOT task. Experiments on three benchmarks, KITTI, nuScenes, and Waymo, show that VPTT achieves state-of-the-art tracking performance on siamese-based tracking running at $\sim$∼62 FPS. Yue Wu 0004, Qiguang Miao, Maoguo Gong, Linghe Kong |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2025 | Triple Point MaskingabstractExisting 3D mask learning methods encounter performance bottlenecks under limited data, and our objective is to overcome this limitation. In this paper, we introduce a triple point masking scheme, named TPM, which serves as a scalable plug-and-play framework for MAE pre-training to achieve multi-mask learning for 3D point clouds. Specifically, we augment the baseline methods with two additional mask choices (i.e., medium mask and low mask) as our core insight is that the recovery process of an object can manifest in diverse ways. Previous high-masking schemes focus on capturing the global representation information but lack fine-grained recovery capabilities, so that the generated pre-training weights tend to play a limited role in the fine-tuning process. With the support of the proposed TPM, current methods can exhibit more flexible and accurate completion capabilities, enabling the potential autoencoder in the pre-training stage to consider multiple representations of a single 3D point cloud object. In addition, during the fine-tuning stage, an SVM-guided weight selection module is proposed to fill the encoder parameters for downstream networks with the optimal weight, maximizing linear accuracy and facilitating the acquisition of intricate representations for new objects. Extensive experimental results and theoretical analysis show that five baselines equipped with the proposed TPM achieve comprehensive performance improvements on various downstream tasks. Our code and models are available athttps://github.com/liujia99/TPM. Linghe Kong, Yue Wu 0004, Maoguo Gong, Hao Li 0009, Qiguang Miao, Wenping Ma 0001, Can Qin |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | Equivariance-Based Markov Decision Process for Unsupervised Point Cloud RegistrationabstractUnsupervised point cloud registration is crucial in 3D computer vision. However, most unsupervised methods struggle to construct effective optimization objectives and reliable unsupervised signals to enhance the performance of the model. To address these issues, with the observation of the significant alignment between the registration process and the Markov Decision Process (MDP), we model point cloud registration as MDP, which can provide more reliable unsupervised signals through the reward. We propose a colored noise based cross-entropy method, which introduces colored noise into sampling process, regulating the power spectral density of the action sequence and expanding the search space, improving the registration effect. Particularly, to strengthen constraints on MDP and training in the transformation space, we utilize equivariance theory to construct transformation equivariant constraint as a new optimization objective and derive equivariant constraint solutions for optimization, providing more reliable unsupervised signals. Extensive experiments demonstrate the superior performance of our method on benchmark datasets. Yue Wu 0004, Jiayi Lei, Yongzhe Yuan, Xiaolong Fan, Maoguo Gong, Wenping Ma 0001, Qiguang Miao, Mingyang Zhang 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2025 | Spatial-Spectral Aggregation Transformer With Diffusion Prior for Hyperspectral Image Super-ResolutionabstractConstrained by imaging systems, hyperspectral images (HSIs) always have a low spatial resolution. Deep learning-based HSI super-resolution methods have achieved impressive results through learning the nonlinear mapping between low-resolution (LR) and high-resolution (HR) images. However, most of them take the LR image or its upsampled version through bicubic interpolation as input, leading to low-quality features and limited details captured by the network. As a powerful generative model, diffusion model has the ability to learn both contextual semantics and textual details from distinct timesteps, enabling the effective exploration of spatial-spectral distributions in high-dimensional data. In this paper, we propose a novel method that extracts high-quality prior information from original images to assist in super-resolution through pretraining a diffusion model. Specifically, we first train a diffusion model using original HSI patches in a self-supervised manner and then obtain prior features from the pretrained denoising U-Net decoder. To efficiently incorporate the prior features into the super-resolution model, we propose an adaptive fusion module based on spatial and spectral attention mechanisms, which enhances features in both dimensions while preserving the original characteristics. Additionally, to leverage the complementarity of spatial and spectral information, we design a spatial-spectral aggregation Transformer module that incorporates an adaptive interaction module to facilitate information exchange across different dimensions, thereby enhancing the representation capability. Extensive experiments on three public hyperspectral datasets demonstrate that the proposed method achieves excellent super-resolution performance and outperforms the state-of-the-art methods in terms of quantitative quality and visual results. Mingyang Zhang 0002, Zhaoyang Wang 0003, Maoguo Gong, Yu Zhou 0051, Fenlong Jiang, Yue Wu 0004 |
IEEE Trans. Circuits Syst. Video Technol. | 8 |
| 2025 | Fast Heterogeneous Multiproblem Surrogates for Transfer Evolutionary Multiobjective OptimizationabstractTransfer evolutionary multiobjective optimization leverages the relevant knowledge from other source problems (distinct but possibly related) to assist the optimization of the target problem of interest. Multi-problem surrogates stack multiple source surrogates to reduce the number of function evaluations of the target expensive problem. The current multi-problem surrogates only considers several source problems and the source and target problems are assumed to be homogeneous. In order to address the above issues, this paper proposes fast heterogeneous multi-problem surrogates for transfer evolutionary multiobjective optimization with a large number of surrogates. First, an iterative surrogate selection strategy is designed to select the highly relevant surrogates from the large-scale surrogate pool to avoid negative transfer. Second, heterogeneous multi-problem surrogates are established to align the features of the source and target models. Finally, an adaptive k-fold cross-validation method is proposed to obtain the predicted values of the target model with low computational costs. Experiments on the multiobjective optimization benchmark problems and multiobjective neural architecture search problems have demonstrated that the proposed method is able to avoid negative transfer in the large-scale scenarios and reduce the computational costs. Hao Li 0009, Pu Xiong, Maoguo Gong, A. K. Qin 0001, Yue Wu 0004, Lining Xing 0001 |
IEEE Trans. Evol. Comput. | 5 |
| 2025 | Multitask Multiscale Feature Selection for Point Cloud Registrationabstract3D point cloud registration is a process of solving the geometric transformation between two point clouds. This process is an important issue in computer vision and pattern recognition. The registration methods based on geometric features are highly sensitive to the scale of feature extraction. Changes in scale can introduce inaccuracies in feature descriptions, thereby compromising the reliability of the registration results. To mitigate the impact of feature scale on the outcomes and the high-dimensional issue arising from features of different scales, we propose a method for multi-scale point cloud feature selection. We solve the high-dimensional problem of feature selection by designing a multi-task framework. By designing a mutual information dimensionality reduction method, we decomposed the high-dimensional feature selection task of different descriptors with multi-scale features into multiple related low-dimensional feature selection tasks. Then, by means of the knowledge transfer among these low-dimensional feature selection tasks, we sought the best feature subset to obtain more robust feature information. We evaluate the effectiveness of our method by conducting extensive experiments on various datasets. The experimental results show that the method outperforms other feature descriptors in terms of descriptive power and robustness and improves the effectiveness of point cloud registration. Yue Wu 0004, Chuang Luo, Maoguo Gong, Hangqi Ding, Jinlong Sheng, Qiguang Miao, Hao Li 0009, Wenping Ma 0001 |
IEEE Trans. Evol. Comput. | 1 |
| 2025 | Evolutionary Multitasking Descriptor Optimization for Point Cloud RegistrationabstractPoint cloud registration (PCR) is an important task for other point cloud tasks. Feature-based methods are widely adopted for their speed and efficiency in PCR. The descriptive capability of features extracted by a single geometric descriptor is limited. Descriptive capabilities can be improved by concatenating features extracted from multiple descriptors. However, due to the existence of redundant and irrelevant features, the correct corresponding points are difficult to match, which further affects the registration effect. We propose an evolutionary multitasking point cloud descriptor optimization method. Integrate existing descriptors to optimize descriptors with stronger description ability. Labeling features to calculate the feature importance for the registration and generating multitasks. In optimized processing, approximate evaluation which is calculated by prior correspondence saved in the database replaces the expensive searching correspondences process in the entire point cloud. Finally, a multiscale filter is developed to remove error correspondences by the geometric information from multiple scale descriptor features. Experimental demonstrate that the proposed approach can optimize a feature subset with higher-descriptive capability compared to other methods and show superior PCR performance on 14 point cloud models. This is the first paper on point cloud descriptor optimization, which provides a new idea for PCR research. Yue Wu 0004, Jinlong Sheng, Hangqi Ding, Peiran Gong, Hao Li 0009, Maoguo Gong, Wenping Ma 0001, Qiguang Miao |
IEEE Trans. Evol. Comput. | 1 |
| 2025 | Physical Adversarial Background Patch Against Aerial Object Detection Based on Pareto EfficiencyabstractFor adversarial attacks on aerial image object detection, some physical background attack methods have been proposed and demonstrated excellent performance. However, most of the methods restrict the protected object to a fixed area, and changes in the position and size of the object can affect the effectiveness of the attack. In addition, the size of the background is only related to the size of the object, without considering any reduction in the background area. To alleviate these issues, this paper proposes a novel adversarial background patch generation strategy, which generates an adversarial background patch where aircrafts parked at any position within it remain undetected. Besides we also aim to make the size of the background as small as possible to improve the concealment of the attack and reduce the economic cost. Specifically, a series of adversarial images are generated by placing the aircraft at random angles and positions on the adversarial background patch. Then, a novel background patch optimization strategy is proposed, which enhances the occlusion robustness of the adversarial background patch by weighting the adversarial loss of each image in the batch based on adversarial difficulty. In addition, this paper proposes a novel area loss for achieving optimal area size, which is used to reduce the cost of producing the patch. Finally, a multi-objective optimization method based on Pareto efficiency is introduced for balancing two conflicting losses, the adversarial loss and the area loss. Experimental results show that the adversarial background patches generated by this method have excellent attack effects in both digital and physical attacks, exhibiting strong occlusion robustness. In addition of that, the adversarial background patches generated by this method achieve effective attacks in a smaller area and reduce the physical implementation cost. Hao Li 0009, Jiachang Li, Maoguo Gong, Haiyue Yu 0001, Kelin Dang, Yu Zhou 0051, A. K. Qin 0001, Yue Wu 0004 |
IEEE Trans. Geosci. Remote. Sens. | 8 |
| 2025 | Dual-Path Prototype Feature Decoupling Alignment Network for Panchromatic and Multispectral ClassificationabstractIn recent years, with the rapid advancements and widespread application of satellite photography technology, it has become increasingly possible to obtain high-quality panchromatic (PAN) and multispectral (MS) data, which has provided new opportunities and challenges for multisource information fusion and classification research. Remote sensing data have the characteristics of small interclass differences and large intraclass differences, which easily leads to category confusion in network learning. In addition, how to fully tap the advantages of multisource data, better align multisource features, improve classification accuracy, and achieve collaborative classification are key issues that need to be solved urgently. In this article, a dual-path prototype feature decoupling alignment network (DPFDA-Net) is designed to solve the above issues. The network consists of two components: a prototype feature embedding (PFE) module and a feature alignment module (FAM) based on prototype decoupling. In the feature extraction stage, the PFE module uses the prototype concept to learn the discriminative prototype features of each category of the dual-source data separately, making the boundaries between categories more obvious. The FAM operates at the dual-source prototype feature level and achieves feature alignment by decoupling single-source prototype features and performing feature transformation to supplement the missing information of another data source. Finally, we use the aligned features for classification. The results of the experiment demonstrate that our approach has made significant progress in improving classification precision. The code is available athttps://github.com/Xidian-AIGroup190726/DPFDANet. Wenping Ma 0001, Yanshan Guo, Hao Zhu 0009, Wenhao Zhao, Mengru Ma, Yue Wu 0004, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2025 | Dense-Weak Ship Detection Based on Foreground-Guided Background Generation Network in SAR ImagesabstractCurrently, ship detection based on Synthetic Aperture Radar (SAR) images still faces significant challenges, particularly in detecting weak and densely distributed ships within complex backgrounds. In areas such as ports and land, the complex background features often resemble those of densely distributed ships, leading to reduced detection accuracy. Additionally, the overlapping and mutual interference of features among dense ships can cause the network to miss detections or produce false positives. Therefore, this paper proposes a Foreground-Guided Background Generation Network (FGBG-Net), which includes a Gaussian Foreground Localization (GFL) model and a Background Feature Removal (BFR) module. The GFL module identifies the approximate high-probability regions of ship foregrounds on the feature map, guiding the network to focus on these regions. The BFR module then progressively removes background interference features based on the positions provided by the GFL module, generating feature maps that are more suitable for detecting weak and dense ships. Our network has been validated on multiple SAR ship datasets, and the experimental results demonstrate noticeable performance improvements, with a mean Average Precision (mAP) increase of 3.4% on the SSDD and HRSID datasets. The relevant code is available at the following link: https://github.com/Xidian-AIGroup190726/FBGBNet/tree/master. Wenping Ma 0001, Xiaoting Yang, Hao Zhu 0009, Xiaoteng Wang, Biao Hou, Mengru Ma, Yue Wu 0004 |
IEEE Trans. Geosci. Remote. Sens. | 7 |
| 2025 | Bidirectional Stacking Ensemble Curriculum Learning for Hyperspectral Image Imbalanced Classification With Noisy LabelsabstractHyperspectral imaging has demonstrated substantial advantages in enhancing classification performance in remote sensing applications due to its abundant spectral information. To address the challenges of label noise and class imbalance in hyperspectral image (HSI) classification, we propose an end-to-end Feature-Guided Network (FGN) for HSI. Instead of merely combining spatial and channel attention, FGN leverages feature-level attention interactions to enhance contextual understanding, leading to better feature extraction, especially for underrepresented classes. Furthermore, a bidirectional loss for curriculum learning (CL) is proposed to rank the HSI training data in a descending or ascending order. The top and bottom loss regularizers are designed to make the proposed model suitable for noisy and imbalanced HSI data distributions. In the phase of selecting pace parameter, a stacking ensemble curriculum learning (SECL) model is established to avoid that the outliers and noisy HSI data are involved into the CL training process. A novel instruction matrix based on sample weights is designed for base classifiers. The outputs of the base models, combined with the expected labels, form the input-output pairs for training the second-level classifier. Experiments conducted on multiple hyperspectral imbalanced datasets with noisy labels demonstrate the superior performance of our method. Yixin Wang 0009, Hao Li 0009, Maoguo Gong, Yue Wu 0004, Peiran Gong, A. K. Qin 0001, Lining Xing 0001, Mingyang Zhang 0002 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2025 | Meta-Collaborative Learning for Arbitrarily Scaled Hyperspectral Image Super-ResolutionabstractDeep learning-based methods for hyperspectral image super-resolution (SR) have achieved significant success in recent years. These methods typically consist of feature extraction module (FEM) and upsampling module. However, due to structural limitations of the upsampling module, most current methods focus on training separate models for different scale factors, which ignores the exploration of potential feature interdependence among different scale factors. In response to these challenges, we introduce a novel framework, called “meta-collaborative learning for arbitrarily scaled hyperspectral image super-resolution” (MCArb). Specifically, MCArb integrates a collaborative learning framework with a meta-learning-based 3-D upsampling module (3DMetaUM) and a scale-aware feature adaptation module (SAFAM). It enables training multiple SR tasks at different scale factors within a single network at the same time. This strategy is able not only to process arbitrary-scale-factor SR for hyperspectral images but also to harness the latent feature interdependence among different scales. In this study, we applied the MCArb framework to transform three deep learning-based hyperspectral image SR networks to MCArb methods, resulting in significant performance enhancements across five hyperspectral datasets. These improvements showcase the proposed MCArb framework’s ability to enhance feature extraction efficiency and capitalize on latent interscale correlations. This code is available athttps://github.com/ShuangWu-XDU/MCArb_HSI_SR. Mingyang Zhang 0002, Maoguo Gong, Fenlong Jiang, Yu Zhou 0051, Yue Wu 0004 |
IEEE Trans. Geosci. Remote. Sens. | 7 |
| 2025 | CCGIB: A Cross-Channel Graph Information Bottleneck PrincipleabstractThe empirical studies of most existing graph neural networks (GNNs) broadly take the original node feature and adjacency relationship as single-channel input, ignoring the rich information of multiple graph channels. To circumvent this issue, the multichannel graph analysis framework has been developed to fuse graph information across channels. How to model and integrate shared (i.e., consistency) and channel-specific (i.e., complementarity) information is a key issue in multichannel graph analysis. In this article, we propose a cross-channel graph information bottleneck (CCGIB) principle to maximize the agreement for common representations and the disagreement for channel-specific representations. Under this principle, we formulate the consistency and complementarity information bottleneck (IB) objectives. To enable optimization, a viable approach involves deriving variational lower bound and variational upper bound (VarUB) of mutual information terms, subsequently focusing on optimizing these variational bounds to find the approximate solutions. However, obtaining the lower bounds of cross-channel mutual information objectives proves challenging through direct utilization of variational approximation, primarily due to the independence of the distributions. To address this challenge, we leverage the inherent property of joint distributions and subsequently derive variational bounds to effectively optimize these information objectives. Extensive experiments on graph benchmark datasets demonstrate the superior effectiveness of the proposed method. Xiaolong Fan, Maoguo Gong, Yue Wu 0004, Mingyang Zhang 0002, Hao Li 0009, Xiangming Jiang |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2025 | One-Nearest Neighborhood Guides Inlier Estimation for Unsupervised Point Cloud RegistrationabstractThe precision of unsupervised point cloud registration methods is typically limited by the lack of reliable inlier estimation and self-supervised signal, especially in partially overlapping scenarios. In this article, we propose an effective inlier estimation method for unsupervised point cloud registration by capturing geometric structure consistency between the source point cloud and its corresponding reference point cloud copy. Specifically, to obtain a high-quality reference point cloud copy, a one-nearest neighborhood (1-NN) point cloud is generated by input point cloud, which facilitates matching map construction and allows for integrating dual neighborhood matching scores of 1-NN point cloud and input point cloud to improve matching confidence. Benefiting from the high-quality reference copy, we argue that the neighborhood graph formed by inlier and its neighborhood should have consistency between source point cloud and its corresponding reference copy. Based on this observation, we construct transformation-invariant geometric structure representations and capture geometric structure consistency to score the inlier confidence for estimated correspondences between source point cloud and its reference copy. This strategy can simultaneously provide the reliable self-supervised signals for model optimization. Finally, we further calculate transformation estimation by the weighted SVD algorithm with the estimated correspondences and the corresponding inlier confidence. We train the proposed model in an unsupervised manner, and extensive experiments on synthetic and real-world datasets illustrate the effectiveness of the proposed method. Yongzhe Yuan, Yue Wu 0004, Maoguo Gong, Qiguang Miao, A. K. Qin 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2025 | Evolutionary Multiobjective Cross-Spectral Adversarial Attacks With Synergistic PatchesabstractDNN have demonstrated vulnerability to adversarial attacks in object detection tasks. While significant progress has been made in single-spectrum attacks, cross-spectral adversarial attacks remain challenging due to the complex tradeoffs between visible and infrared domains. To address this, an evolutionary multiobjective cross-spectral attack (MoXAttack) framework, for developing adversarial patches in closed-box cross-spectral scenarios is proposed. MoXAttack incorporates a multipopulation constraint-handling technique, which uses both penalty functions and feasibility rules to guide the search process. Spectrum-aware genetic operators are introduced to enhance solution diversity and feasibility. The framework automatically optimizes the smooth to cross-spectral shared patch shape using curvature energy. In addition, MoXAttack utilizes SVD for visible spectrum texture perturbations and adjustable thermal shielding material thickness for infrared spectrum control. Experiments on the LLVIP dataset demonstrate that MoXAttack achieves competitive performance across multiple object detection models. Ablation studies reveal the positive impact of improved components on attack effectiveness. The multipatch strategy improves attack success rates by at least 17%, while optimized patch shapes outperform conventional geometric shapes by at least 25% in terms of mAP drop. In the physical world test, the proposed method shows stability in different viewing angles. Wencheng Han, Hao Li 0009, Maoguo Gong, Yue Wu 0004, A. K. Qin 0001, Lining Xing 0001, Yu Zhou 0051 |
IEEE Trans. Syst. Man Cybern. Syst. | 4 |
| 2025 | Divide-and-Conquer Evolutionary Multitasking OptimizationabstractThis article proposes a novel evolutionary multitasking optimization (EMTO) paradigm called divide-and-conquer EMTO, which divides the original complex optimization problem into multiple simple optimization tasks and then these tasks are optimized by EMTO concurrently to formulate the resulting solution of the original problem. The main characteristics of divide-and-conquer EMTO are that the considered problem can be divided into multiple small-scale optimization tasks and the optimal solution is the combination of solutions of all tasks. In order to achieve the optimal combined fitness of all tasks, a relative improvement function and an adaptive exploration optimization strategy are designed for dynamic resource allocation across tasks. Finally, a case study on hyperspectral unmixing is investigated in the proposed divide-and-conquer EMTO framework by dividing the hyperspectral image into several homogeneous regions to formulate multiple sparse unmixing tasks. Experiments on benchmark and sparse unmixing problems demonstrate the superiority of divide-and-conquer EMTO. Hao Li 0009, Maoguo Gong, Yue Wu 0004, A. K. Qin 0001, Lining Xing 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 4 |
| 2025 | Pseudo Label Learning for Partial Point Cloud RegistrationabstractPartial point cloud registration plays a crucial role in computer vision and has widespread applications in 3D map construction, pose estimation, and high-precision localization. However, the collected point clouds often contain missing data due to hardware limitations and complex environments. Various partial registration algorithms have been proposed, most of which rely on estimating overlap regions. However, a significant proportion of these algorithms rely heavily on ground truth labels. Manual labeling is both time-consuming and labor-intensive, whereas algorithmic automatic labeling lacks sufficient accuracy. To tackle this issue, we present PSEudo Label learning for unsupervised partial point cloud registration (PSEL). This method utilizes complementary tasks to learn reliable pseudo labels for overlap regions and correspondences without depending on ground truth labels. The key idea is to use the complementarity between overlap estimation and registration to generate two types of pseudo labels based on the nearest points in pairs of aligned point clouds. These pseudo labels are then employed to supervise the learning of overlap regions and correspondences, gradually enhancing their accuracy throughout the learning process and ultimately establishing an unsupervised learning framework. PSEL consists of an overlap estimation module and a correspondence filtering module. The pseudo labels generated after registration are used to supervise both modules. Notably, the correspondence filtering module has two pipelines. The similarity and difference of the corresponding point features are used to eliminate false correspondences during the training and inference stages, respectively, with only the latter being optimized with pseudo labels. To validate the effectiveness of our registration method, we conducted experiments using the synthetic dataset ModelNet40, the indoor dataset 3DMatch, and the outdoor dataset KITTI. Wenping Ma 0001, Yue Wu 0004, Yue Zhang 0040, Hao Zhu 0009, Biao Hou, Licheng Jiao |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2024 | Neural Gaussian Similarity Modeling for Differential Graph Structure LearningabstractGraph Structure Learning (GSL) has demonstrated considerable potential in the analysis of graph-unknown non-Euclidean data across a wide range of domains. However, constructing an end-to-end graph structure learning model poses a challenge due to the impediment of gradient flow caused by the nearest neighbor sampling strategy. In this paper, we construct a differential graph structure learning model by replacing the non-differentiable nearest neighbor sampling with a differentiable sampling using the reparameterization trick. Under this framework, we argue that the act of sampling nearest neighbors may not invariably be essential, particularly in instances where node features exhibit a significant degree of similarity. To alleviate this issue, the bell-shaped Gaussian Similarity (GauSim) modeling is proposed to sample non-nearest neighbors. To adaptively model the similarity, we further propose Neural Gaussian Similarity (NeuralGauSim) with learnable parameters featuring flexible sampling behaviors. In addition, we develop a scalable method by transferring the large-scale graph to the transition graph to significantly reduce the complexity. Experimental results demonstrate the effectiveness of the proposed methods. Xiaolong Fan, Maoguo Gong, Yue Wu 0004, Zedong Tang, Jieyi Liu |
AAAI | 3 |
| 2024 | M3SOT: Multi-Frame, Multi-Field, Multi-Space 3D Single Object Trackingabstract3D Single Object Tracking (SOT) stands a forefront task of computer vision, proving essential for applications like autonomous driving. Sparse and occluded data in scene point clouds introduce variations in the appearance of tracked objects, adding complexity to the task. In this research, we unveil M3SOT, a novel 3D SOT framework, which synergizes multiple input frames (template sets), multiple receptive fields (continuous contexts), and multiple solution spaces (distinct tasks) in ONE model. Remarkably, M3SOT pioneers in modeling temporality, contexts, and tasks directly from point clouds, revisiting a perspective on the key factors influencing SOT. To this end, we design a transformer-based network centered on point cloud targets in the search area, aggregating diverse contextual representations and propagating target cues by employing historical frames. As M3SOT spans varied processing perspectives, we've streamlined the network—trimming its depth and optimizing its structure—to ensure a lightweight and efficient deployment for SOT applications. We posit that, backed by practical construction, M3SOT sidesteps the need for complex frameworks and auxiliary components to deliver sterling results. Extensive experiments on benchmarks such as KITTI, nuScenes, and Waymo Open Dataset demonstrate that M3SOT achieves state-of-the-art performance at 38 FPS. Our code and models are available at https://github.com/ywu0912/TeamCode.git. Yue Wu 0004, Maoguo Gong, Qiguang Miao, Wenping Ma 0001, Can Qin |
AAAI | 2 |
| 2024 | Reliable Conflictive Multi-View LearningabstractMulti-view learning aims to combine multiple features to achieve more comprehensive descriptions of data. Most previous works assume that multiple views are strictly aligned. However, real-world multi-view data may contain low-quality conflictive instances, which show conflictive information in different views. Previous methods for this problem mainly focus on eliminating the conflictive data instances by removing them or replacing conflictive views. Nevertheless, real-world applications usually require making decisions for conflictive instances rather than only eliminating them. To solve this, we point out a new Reliable Conflictive Multi-view Learning (RCML) problem, which requires the model to provide decision results and attached reliabilities for conflictive multi-view data. We develop an Evidential Conflictive Multi-view Learning (ECML) method for this problem. ECML first learns view-specific evidence, which could be termed as the amount of support to each category collected from data. Then, we can construct view-specific opinions consisting of decision results and reliability. In the multi-view fusion stage, we propose a conflictive opinion aggregation strategy and theoretically prove this strategy can exactly model the relation of multi-view common and view-specific reliabilities. Experiments performed on 6 datasets verify the effectiveness of ECML. The code is released at https://github.com/jiajunsi/RCML. Jiajun Si, Ziyu Guan, Wei Zhao 0019, Yue Wu 0004, Xiyue Gao |
AAAI | 5 |
| 2024 | Evolutionary Multitasking with Compatibility Graph for Point Cloud Registrationabstract3D point cloud registration is a fundamental task in computer vision, aimed at estimating a transformation to align a pair of point clouds. For point cloud registration, where the popular methods are used to build a compatibility graph. Different methods of constructing compatibility graphs can result in different information regions to be searched in the graph, due to varying constraints. Evolutionary multitask optimization has gained attention in the field of evolutionary computation, as it enables knowledge transfer among multitasks to enhance the exploration of information. Inspired by this theory, this paper proposes a method to utilize compatibility graphs as tasks through evolutionary multitasking for solving the problem of point cloud registration. We first construct two compatibility graph tasks with different tightness constraints for point cloud registration. Due to the different constraints of the two proposed tasks, the local information of the graph of interest might be biased. This bias can be effectively utilized in the evolutionary multitasking framework to enhance the ability to discover meaningful consensus relationships in the search space of the graph. Then we map the two tasks to a unified search space by designing clique selection strategies, and carrying out knowledge transfer between the two tasks, aiming to emphasize more on the local consensus information in the graphs. Lastly, the efficacy of our proposed method is further validated on multiple registration datasets. Hangqi Ding, Yue Wu 0004, Hao Li 0009, Maoguo Gong, Wenping Ma 0001, Qiguang Miao |
CEC | 3 |
| 2024 | One-Shot Surrogate for Evolutionary Multiobjective Neural Architecture SearchabstractNovel benchmarks for multiobjective neural archi-tecture search are emerging consistently. It stimulates the need of knowledge transfer techniques to facilitate the repetitive and cumbersome search process. One-shot surrogate is hereby proposed to transfer knowledge from source problems. Specif-ically, a Pareto-aware super-surrogate construction technique is proposed aiming at efficient exploitation of knowledge from benchmarks to make transfer easier. Then, during the surrogate prediction process, appropriate sub-surrogates are sampled from the super-surrogate to jointly guide the evaluation decision. The multiobjective surrogate is also redesigned as objective-wise, decision-making, and combinative ones, so that it can make transfer flexible, as well as evade from the scalarization issue in performance measurement. This knowledge transfer scheme assists the convergence of the search process using previously constructed surrogates from multiple source problems, balancing exploration and exploitation. Kuangda Lyu, Maoguo Gong, Hao Li 0009, Yuan Gao 0019, Yue Wu 0004, Dan Feng 0002, Jiao Shi, Yu Lei 0002 |
CEC | 5 |
| 2024 | Inlier Confidence Calibration for Point Cloud RegistrationabstractInliers estimation constitutes a pivotal step in partially overlapping point cloud registration. Existing methods broadly obey coordinate-based scheme, where inlier con-fidence is scored through simply capturing coordinate differences in the context. However, this scheme results in massive inlier misinterpretation readily, consequently affecting the registration performance. In this paper, we explore to extend a new definition called inlier confidence calibration (ICC) to alleviate the above issues. Firstly, we provide finely initial correspondences for ICC in order to generate high quality reference point cloud copy corresponding to the source point cloud. In particular, we develop a soft assignment matrix optimization theorem that offers faster speed and greater precision compared to Sinkhorn. Benefiting from the high quality reference copy, we argue the neighborhood patch formed by inlier and its neighborhood should have consistency between source point cloud and its reference copy. Based on this insight, we construct transformation-invariant geometric constraints and capture geometric structure consistency to calibrate inlier confidence for estimated correspondences between source point cloud and its reference copy. Finally, transformation is further calculated by the weighted SVD algorithm with the calibrated inlier confidence. Our model is trained in an unsupervised manner, and extensive experiments on synthetic and real-world datasets illustrate the effectiveness of the proposed method. Yongzhe Yuan, Yue Wu 0004, Xiaolong Fan, Maoguo Gong, Qiguang Miao, Wenping Ma 0001 |
CVPR | 2 |
| 2024 | Evolutionary Multitasking with Two-level Knowledge Transfer for Multi-view Point Cloud RegistrationabstractPoint cloud registration is a hot research topic in the field of computer vision. In recent years, the registration method based on evolutionary computation has attracted more and more attention because of its robustness to initial pose and flexibility of objective function design. However, most of the current evolutionary computation-based point cloud registration methods do not take into account the multi-view problem, that is, to capture the close relationship between point clouds from different perspectives. We fully realize that if these relations are used correctly, the registration performance can be improved. Therefore, this paper proposes an evolutionary multitasking multi-view point cloud registration method, which solves the problem of multi-view error accumulation. To ensure the unity of global and local, a two-level knowledge transfer strategy is proposed, which divides the multi-view cloud registration task into two levels. This strategy unifies the search space of two registration tasks, solves the negative transfer phenomenon, and avoids the problem of falling into the local optimum. Finally, the effectiveness of the method is verified by sufficient experiments. This method has strong robustness to noise and outliers, and can be effectively implemented in various registration scenarios. Hangqi Ding, Haoran Xu 0005, Yue Wu 0004, Hao Li 0009, Maoguo Gong, Wenping Ma 0001, Qiguang Miao, Jiao Shi, Yu Lei 0002 |
GECCO | 3 |
| 2024 | PointMC: Multi-instance Point Cloud Registration based on Maximal CliquesabstractMulti-instance point cloud registration is the problem of estimating multiple rigid transformations between two point clouds. Existing solutions rely on global spatial consistency of ambiguity and the time-consuming clustering of highdimensional correspondence features, making it difficult to handle registration scenarios where multiple instances overlap. To address these problems, we propose a maximal clique based multiinstance point cloud registration framework called PointMC. The key idea is to search for maximal cliques on the correspondence compatibility graph to estimate multiple transformations, and cluster these transformations into clusters corresponding to different instances to efficiently and accurately estimate all poses. PointMC leverages a correspondence embedding module that relies on local spatial consistency to effectively eliminate outliers, and the extracted discriminative features empower the network to circumvent missed pose detection in scenarios involving multiple overlapping instances. We conduct comprehensive experiments on both synthetic and real-world datasets, and the results show that the proposed PointMC yields remarkable performance improvements. Yue Wu 0004, Xidao Hu, Yongzhe Yuan, Xiaolong Fan, Maoguo Gong, Hao Li 0009, Mingyang Zhang 0002, Qiguang Miao, Wenping Ma 0001 |
ICML | 1 |
| 2024 | MCDet: Multi-Content Collaboration Detector for Multiscale Remote Sensing ObjectabstractIn previous works, powerful CNN backbones are typically used for one- or two-stage detectors to facilitate multi-categories object classification. Unfortunately, continuous convolution and pooling operations tend to weaken the detailed information. We propose an end-to-end Multi-content Collaboration Detector (MCDet) to improve object recognition accuracy. First, we summarize the reasons for the disappearance of detailed features in traditional feature extraction backbone networks, and propose a Shallow Clue Refinement (SCR) module, which helps us to retain more critical local detail information in the downsampling process. Second, to receive more suitable contextual information, we design a Self-dilating Spatial Pooling (SSP) module, it adaptively learns a contextual reception field, thereby alleviating the mismatch between the theoretical receptive field of the network design and the practical requirements. Finally, extensive experiments on the NWPU VHR-10 and DIOR datasets have shown that the proposed MCDet significantly improves detection accuracy. Our code is available at https://github.com/Xidian-AIGroup190726/RS-objectdetection-MCDet. Wenping Ma 0001, Hao Zhu 0009, Yue Wu 0004, Biao Hou, Licheng Jiao |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2024 | ShiftAttack: Toward Attacking the Localization Ability of Object DetectorabstractState-of-the-art (SOTA) adversarial attacks expose vulnerabilities in object detectors, often resulting in erroneous predictions. However, existing adversarial attacks neglect the stealth and flexibility of adversarial examples, which are crucial for conducting contextually consistent and inconspicuous attacks. To address these issues, leveraging the observed phenomenon of predicted box offsets in real-world object detection scenarios, this paper presents a novel adversarial attack framework called ShiftAttack. It leverages the concept of dense detection in prevalent object detectors, by boosting the confidence of low Intersection over Union (IoU) predictions within the positive samples (the set of predicted boxes responsible for localizing the same target), which leads to the erroneous exclusion of true positive predictions during the post-processing stage. Such a paradigm is highly stealthy as the shifted predictions seem like natural detector mistakes rather than obvious manipulations. To enhance the flexibility of ShiftAttack this paper proposes a generative approach called ShiftAttack Generator (SAG), which can not only shift predicted boxes for any target in arbitrary directions and distances but also facilitate adaptive feature exchange between pre- and post-shift regions to optimize the attack. Additionally, the proposed SAG incorporates the Dynamic Hinge Loss (DHL) to ensure the imperceptibility of perturbations, effectively mitigating the Patch-Pattern associated with the use of$\mathcal {L}_{2}$norm. Extensive experiments confirm that SAG surpasses other SOTA adversarial attacks in effectiveness, speed and stealthiness. Hao Li 0009, Maoguo Gong, Shiguo Chen, A. K. Qin 0001, Zhenxing Niu, Yue Wu 0004, Yu Zhou 0051 |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2024 | Learning Discriminative Features via Multi-Hierarchical Mutual Information for Unsupervised Point Cloud RegistrationabstractExtracting discriminative representations is the key step for correspondence-free point cloud registration. The extracted representations require to be discriminative to transformation, which demands representations to reduce the influence of redundant information irrelevant to transformation. However, recently proposed methods ignore this crucial property, resulting in limited ability to represent point cloud. In addition, researching correspondence-free point cloud registration has stagnated in recent years. In this paper, we try to relieve features redundancy issue for correspondence-free point cloud registration from a new perspective. Specifically, our method comprises two stages: feature extraction stage and rigid body transformation stage. In feature extraction stage, we aim to maximize multi-hierarchical mutual information between different hierarchical features, which can provide discriminative and less redundancy representations to regress transformation parameters for next stage. In rigid body transformation stage, we utilize dual quaternion to estimate transformation parameters, which combines rotation and translation simultaneously within a unified framework and obtains a compact representations for rigid transformation. The proposed model is trained in an unsupervised manner on the ModelNet40 dataset. The experimental results illustrate that our method achieves higher accuracy and robustness compared with existing correspondence-free methods. Yongzhe Yuan, Yue Wu 0004, Mingyu Yue, Maoguo Gong, Xiaolong Fan, Wenping Ma 0001, Qiguang Miao |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | Privacy-Enhanced Multitasking Particle Swarm Optimization Based on Homomorphic EncryptionabstractEvolutionary multitasking optimization (EMTO) is a new optimization paradigm proposed in the field of evolutionary computation in recent years. EMTO can solve several different optimization tasks simultaneously and facilitate superior convergence characteristics by transferring effective knowledge among the tasks. However, existing EMTO usually focuses only on facilitating convergence characteristics while neglecting the potential privacy leakage problem in the knowledge transfer between different tasks. The privacy leakage could result in considerable financial losses or severe reputation impairment, which may impede the development of EMTO in real-world applications. To solve the problem of privacy leakage in EMTO, this paper proposes a privacy-enhanced multitasking particle swarm optimization algorithm. A knowledge transfer strategy with privacy preservation is designed based on homomorphic encryption by combining multitasking particle swarm optimization. In addition, an inter-task knowledge transfer mechanism implemented in a low-dimensional subspace is introduced to reduce the extra computational burden caused by privacy preservation. Comprehensive experiments are conducted on synthetic and NAS problems to verify the effectiveness of the proposed method. According to the experimental results, the proposed method has remarkable advantages in privacy preservation compared to existing EMTO. Hao Li 0009, Fanggao Wan, Maoguo Gong, A. K. Qin 0001, Yue Wu 0004, Lining Xing 0001 |
IEEE Trans. Evol. Comput. | 5 |
| 2024 | Evolutionary Multiform Optimization With Two-Stage Bidirectional Knowledge Transfer Strategy for Point Cloud RegistrationabstractPoint cloud registration is an important task in computer vision, where the goal is to estimate a transformation to align a pair of point clouds. Most of the existing registration methods face the problems of poor robustness and getting stuck in local optima. Evolutionary multitasking is an effective paradigm to enhance global search capability and improve convergence characteristics through knowledge transfer across multiple related tasks. Inspired by evolutionary multitasking, this article proposes a multiform optimization approach through evolutionary multitasking for solving the point cloud registration problems. We first construct two related registration tasks with different functional landscapes to form a multiform optimization problem. Compared with methods that only focus on a single registration attribute, the two proposed tasks focus on robustness and precision of registration, respectively. Then, a new two-stage bidirectional knowledge transfer strategy is presented, which can implement efficient knowledge transfer among two related tasks. Finally, both simulations and real experiments show the power of our method. The proposed method is robust to noise, outliers, and partial overlaps and is effective in multiple real registration scenarios, such as object registration, scene reconstruction, and simultaneous localization and mapping. Yue Wu 0004, Hangqi Ding, Maoguo Gong, A. K. Qin 0001, Wenping Ma 0001, Qiguang Miao, Kay Chen Tan |
IEEE Trans. Evol. Comput. | 1 |
| 2024 | Dual Appearance-Aware Enhancement for Oriented Object DetectionabstractOriented detectors have become the mainstream of object detection in remote-sensing images since they provide more precise bounding boxes and contain less background. However, there remain several challenges that restrict the detection performance and need to be tackled. This article focuses on the following two aspects: 1) numerous tiny objects in remote-sensing images pose a challenge for the detectors pursuing high recall and accurate localization and 2) specific categories with large aspect ratios and arbitrary angles also trouble the regression of the detectors. We attempt to alleviate the above problems by constructing a weak feature extraction network (WFEN) and a dual appearance-aware (DA) loss. Specifically, WFEN is used to extract hierarchical weight vectors for multiscale feature layers by employing a lightweight convolutional module, aiming to fuse activation features distributed in different layers and provide pure features for subsequent regression and classification. DA loss is tailored to regressions of tiny and slender objects by dynamically modulating the associated loss on objects with various appearances, which consists of two auxiliary losses, termed scale-aware loss${\mathcal {L}}_{S}$and aspect-ratio-aware loss${\mathcal {L}}_{A}$. These two components can contribute to each other, that is, the former provides more accurate features for detection tasks, while the latter can reciprocate the former by imposing constraints on crucial objects, and together constitute an appearance sensitivity detector (ASDet). Extensive experiments on three public datasets demonstrate that our ASDet outperforms all refine-stage detectors in terms of accuracy while maintaining the superior inference speed of single-stage counterparts. Maoguo Gong, Hongyu Zhao 0007, Yue Wu 0004, Zedong Tang, Kaiyuan Feng |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | MSANet: Multiscale Self-Attention Aggregation Network for Few-Shot Aerial Imagery SegmentationabstractFew-shot aerial imagery segmentation refers to the task of segmenting specific objects in scenes that have not been encountered during training with a small amount of annotated data for reference. However, most existing few-shot segmentation algorithms are primarily designed for natural images, and there is still a lack of exploration in the context of remote sensing aerial imagery. In this article, we propose a novel multiscale self-attention aggregation network (MS2A2Net), dubbed MS2A2Net, to address the challenge of few-shot aerial image segmentation in terms of scarce data and network architecture. Specifically, we first incorporate the designed asymmetric momentum contrastive learning (AMCL) into the pre-training stage, to improve the representation capability of the backbone without the expensive labeled data. Then the frozen encoder is transferred to the downstream few-shot segmentation task as the feature embedding. In terms of network architecture, we design self-attention aggregation in multiscale feature fusion, to construct the dual correlation of foreground and background between support and query features at the pixel level. Besides, the coordinate attention is designed to rearrange the distribution of feature importance in both horizontal and vertical spatial order perspectives, which facilitates adaptive fusion with the multiscale features. To verify the availability of the proposed MS2A2Net, we also reconstructed two novel datasets dedicated to few-shot aerial image segmentation, called DLRSD-$4^{i}$and iSAID-$4^{i}$. The experimental results show that our approach MS2A2Net is superior in three few-shot benchmark aerial imagery segmentation datasets, which achieves competitive segmentation performance. Extensive ablation experiments also reflect the effectiveness and scalability of the proposed components and overall network architecture. Jianzhao Li, Maoguo Gong, Mingyang Zhang 0002, Yourun Zhang, Shanfeng Wang, Yue Wu 0004 |
IEEE Trans. Geosci. Remote. Sens. | 7 |
| 2024 | Robust Self-Paced Incremental Learning for Multitemporal Remote Sensing Image ClassificationabstractClassification of multitemporal remote sensing (MTRS) images with only a few labels of one of these images has attracted widespread interest in recent years. It usually confronts three problems: domain increment, class increment, and class disappearance. In this article, a robust self-paced incremental learning (RSPIL) is proposed to alleviate the above problems. First, change detection is used to transfer labels from the source image to the target one for formulating a combined training set. Then, the weighted classification loss and the distillation loss are considered to ensure classification performance and minimal forgetting. In particular, a novel entropy-inhibit loss is proposed to suppress the classification capability for the disappearing classes. These losses are combined with self-paced learning (SPL) by introducing a weight variable to measure the “easiness” of the training samples, which is able to automatically acquire accurate decision boundaries from easy to hard since the combined training set generated by change detection contains noisy samples and outliers. Finally, a nearest-average-eigenvector classifier and an exemplar set management strategy based on the sample weights are designed to alleviate catastrophic forgetting (CF). Twenty-four MTRS image datasets from four areas are considered in the experiments. The classification results demonstrate that the proposed method is able to alleviate CF and achieves significant improvements on multitemporal image datasets. Hao Li 0009, Pengyang Niu, Maoguo Gong, Lining Xing 0001, Yue Wu 0004, A. K. Qin 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2024 | Intra- and Intersource Interactive Representation Learning Network for Remote Sensing Images ClassificationabstractRecently, remote sensing technology has developed faster and faster, and obtaining high-quality panchromatic (PAN) and multispectral (MS) images has become more accessible. The complementarity between them provides new opportunities in multisource remote sensing image classification. However, solving the problem of the semantic gap between multisource high-level features and, at the same time, utilizing the complementary properties between them to reduce intersource information redundancy is still a challenge. This article constructs an$I^{3}$RL-Net for the multisource remote sensing image classification task. Specifically, we design a cross-source interactive enhanced fusion module (CIEF-Module). For multilevel multisource features, by strengthening the dependencies of intrasource features and conducting intersource enhanced fusion, intrasource correlation features are refined, and the problem of the intersource semantic gap can be effectively alleviated. During the cross-source interaction process, we design a complementary representation supervised learning strategy (CRSL-Strategy). According to the similarities and differences of multisource features, it can adaptively promote complementary feature learning, thus generating a nonredundant multisource representation. The method has been verified to be effective on multiple RS datasets. The code is open source at:https://github.com/Xidian-AIGroup190726/Ping-Pie-I3RL-Net.git. Wenping Ma 0001, Yanshan Guo, Hao Zhu 0009, Xiaoyu Yi 0002, Wenhao Zhao, Yue Wu 0004, Biao Hou, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2024 | Adaptive Feature Separation Network for Remote Sensing Object DetectionabstractWith the development of remote sensing technology, remote sensing object detection has been widely applied in various fields, but it still faces some thorny challenges, such as the following: 1) the complexity of object scale changes in remote sensing images makes it difficult to improve the performance of small object detection and 2) remote sensing images have complex backgrounds and densely arranged small and weak objects, which pose a serious problem of feature interference. To alleviate these challenges, we propose an end-to-end adaptive feature separation network called AFSNet, which includes a scale-aware module (SAM) and a class-aware module (CAM). The SAM mainly enables feature maps of different resolutions to detect objects of different scales. Shallow feature maps mainly suppress the features of large objects they contain to focus on small object detection, while deep feature maps increase the detailed features of large objects they contain to focus on large object detection. The CAM is mainly used to distinguish the features in the feature map by category, separating the features of different categories into different channels, thus mitigating the problem of inter class feature interference, and blocking background interference. The effectiveness of this article has been proven on the NWPU VHR-10, IPIU-M, DIOR, and DOTA2.0 datasets. It can be widely applied in civilian, military, and other fields. Through experimental verification, our AFSNet achieved 97.70% mAP on the NWPU VHR-10 dataset, 78.9% mAP on the DIOR dataset, and 58.22% mAP on the DOTA2.0 dataset. Our code is available at:https://github.com/Xidian-AIGroup190726/AFSNet. Wenping Ma 0001, Yiting Wu, Hao Zhu 0009, Wenhao Zhao, Yue Wu 0004, Biao Hou, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2024 | A Multi-Modal Vertical Federated Learning Framework Based on Homomorphic EncryptionabstractFederated learning has gained prominence as an effective solution for addressing data silos, enabling collaboration among multiple parties without sharing their data. However, existing federated learning algorithms often neglect the challenge posed by multi-modal data distribution. Moreover, previous pioneering work face limitations in encrypting the exponential and logarithmic operations of the objective function with multiple independent variables, and they rely on a third-party cooperator for encryption. To address these limitations, this paper introduces a universal multi-modal vertical federated learning framework. To tackle the data distribution challenge, we propose a two-step multi-modal transformer model that captures cross-domain semantic features effectively. For encryption, where traditional additively homomorphic encryption algorithms fall short by supporting only addition and multiplication, we employ bivariate Taylor series expansion to transform the objective function. Integrating these components, we present a comprehensive training and transmission protocol that eliminates the need for a third-party cooperator during the encryption process. Extensive experiments conducted on diverse video-text and image-text datasets validate the superior performance of our framework compared to state-of-the-art approaches, affirming its effectiveness in multi-modal vertical federated learning settings. Maoguo Gong, Yuanqiao Zhang, Yuan Gao 0019, A. K. Qin 0001, Yue Wu 0004, Shanfeng Wang, Yihong Zhang 0008 |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2024 | Inter-Modal Masked Autoencoder for Self-Supervised Learning on Point CloudsabstractMasked autoencoder (MAE) is a recently widely used self-supervised learning method that has achieved great success in NLP and computer vision. However, the potential advantages of masked pre-training for point cloud understanding have not been fully explored. There is preliminary work on MAE-based point clouds using the Transformer architecture to explore low-level geometric representations in 3D space, which is insufficient for fine-grained decoding completion and downstream tasks. Inspired by multimodality, we propose Inter-MAE, a inter-modal MAE method for self-supervised learning on point clouds. Specifically, we first use Point-MAE as a baseline to partition point clouds into random low percentage of visible and high percentage of masked point patches. Then, a standard Transformer-based autoencoder is built by asymmetric design and shifting mask operations, and latent features are learned from the visible point patches aiming to recover the masked point patches. In addition, we generate image features based on ViT after point cloud rendering to form inter-modal contrastive learning with the decoded features of the completed point patches. Extensive experiments show that the proposed Inter-MAE generates pre-trained models that are effective and exhibit superior results in various downstream tasks. For example, an accuracy of 85.4% is achieved on ScanObjectNN and 86.3% on ShapeNetPart, outperforming other state-of-the-art self-supervised learning methods. Notably, our work establishes for the first time the feasibility of applying image modality to masked point clouds. Yue Wu 0004, Maoguo Gong, Zhixiao Liu, Qiguang Miao, Wenping Ma 0001 |
IEEE Trans. Multim. | 2 |
| 2024 | Self-Supervised Intra-Modal and Cross-Modal Contrastive Learning for Point Cloud UnderstandingabstractLearning effective representations from unlabeled data is a challenging task for point cloud understanding. As the human visual system can map concepts learned from 2D images to the 3D world, and inspired by recent multimodal research, we introduce data from point cloud modality and image modality for joint learning. Based on the properties of point clouds and images, we propose CrossNet, a comprehensive intra- and cross-modal contrastive learning method that learns 3D point cloud representations. The proposed method achieves 3D-3D and 3D-2D correspondences of objectives by maximizing the consistency of point clouds and their augmented versions, and with the corresponding rendered images in invariant space. We further distinguish the rendered images into RGB and grayscale images to extract color and geometric features, respectively. These training objectives combine feature correspondences between modalities to combine rich learning signals from point clouds and images. Our CrossNet is simple: we add a feature extraction module and a projection head module to the point cloud and image branches, respectively, to train the backbone network in a self-supervised manner. After the network is pretrained, only the point cloud feature extraction module is required for fine-tuning and directly predicting results for downstream tasks. Our experiments on multiple benchmarks demonstrate improved point cloud classification and segmentation results, and the learned representations can be generalized across domains. Yue Wu 0004, Maoguo Gong, Peiran Gong, Xiaolong Fan, A. K. Qin 0001, Qiguang Miao, Wenping Ma 0001 |
IEEE Trans. Multim. | 1 |
| 2024 | MPCT: Multiscale Point Cloud Transformer With a Residual NetworkabstractThe self-attention (SA) network revisits the essence of data and has achieved remarkable results in text processing and image analysis. SA is conceptualized as a set operator that is insensitive to the order and number of data, making it suitable for point sets embedded in 3D space. However, working with point clouds still poses challenges. To tackle the issue of exponential growth in complexity and singularity induced by the original SA network without position encoding, we modify the attention mechanism by incorporating position encoding to make it linear, thus reducing its computational cost and memory usage and making it more feasible for point clouds. This article presents a new framework called multiscale point cloud transformer (MPCT), which improves upon prior methods in cross-domain applications. The utilization of multiple embeddings enables the complete capture of the remote and local contextual connections within point clouds, as determined by our proposed attention mechanism. Additionally, we use a residual network to facilitate the fusion of multiscale features, allowing MPCT to better comprehend the representations of point clouds at each stage of attention. Experiments conducted on several datasets demonstrate that MPCT outperforms the existing methods, such as achieving accuracies of 94.2% and 84.9% in classification tasks implemented on ModelNet40 and ScanObjectNN, respectively. Yue Wu 0004, Maoguo Gong, Zhixiao Liu, Qiguang Miao, Wenping Ma 0001 |
IEEE Trans. Multim. | 1 |
| 2024 | A Collaborative Multimodal Learning-Based Framework for COVID-19 DiagnosisabstractThe pandemic of coronavirus disease 2019 (COVID-19) has led to a global public health crisis, which caused millions of deaths and billions of infections, greatly increasing the pressure on medical resources. With the continuous emergence of viral mutations, developing automated tools for COVID-19 diagnosis is highly desired to assist the clinical diagnosis and reduce the tedious workload of image interpretation. However, medical images in a single site are usually of a limited amount or weakly labeled, while integrating data scattered around different institutions to build effective models is not allowed due to data policy restrictions. In this article, we propose a novel privacy-preserving cross-site framework for COVID-19 diagnosis with multimodal data, seeking to effectively leverage heterogeneous data from multiple parties while preserving patients' privacy. Specifically, a Siamese branched network is introduced as the backbone to capture inherent relationships across heterogeneous samples. The redesigned network is capable of handling semisupervised inputs in multimodalities and conducting task-specific training, in order to improve the model performance of various scenarios. The framework achieves significant improvement compared with state-of-the-art methods, as we demonstrate through extensive simulations on real-world datasets. Yuan Gao 0019, Maoguo Gong, Yew-Soon Ong, A. K. Qin 0001, Yue Wu 0004, Fei Xie 0007 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2024 | Explore the Influence of Shallow Information on Point Cloud RegistrationabstractFeature extraction is a key step for deep-learning-based point cloud registration. In the correspondence-free point cloud registration task, the previous work commonly aggregates deep information for global feature extraction and numerous shallow information which is positive to point cloud registration will be ignored with the deepening of the neural network. Shallow information tends to represent the structural information of the point cloud, while deep information tends to represent the semantic information of the point cloud. In addition, fusing information of different dimensions is conducive to making full use of shallow information. Inspired by this, we verify shallow information in the middle layers can bring a positive impact on the point cloud registration task. We design various architectures to combine shallow information and deep information to extract global features for point cloud registration. Experimental results on the ModelNet40 dataset illustrate that feature extractors that incorporate shallow information will bring positive performance. Wenping Ma 0001, Mingyu Yue, Yue Wu 0004, Yongzhe Yuan, Hao Zhu 0009, Biao Hou, Licheng Jiao |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2024 | RORNet: Partial-to-Partial Registration Network With Reliable Overlapping RepresentationsabstractThree-dimensional point cloud registration is an important field in computer vision. Recently, due to the increasingly complex scenes and incomplete observations, many partial-overlap registration methods based on overlap estimation have been proposed. These methods heavily rely on the extracted overlapping regions with their performances greatly degraded when the overlapping region extraction underperforms. To solve this problem, we propose a partial-to-partial registration network (RORNet) to find reliable overlapping representations from the partially overlapping point clouds and use these representations for registration. The idea is to select a small number of key points called reliable overlapping representations from the estimated overlapping points, reducing the side effect of overlap estimation errors on registration. Although it may filter out some inliers, the inclusion of outliers has a much bigger influence than the omission of inliers on the registration task. The RORNet is composed of overlapping points' estimation module and representations' generation module. Different from the previous methods of direct registration after extraction of overlapping areas, RORNet adds the step of extracting reliable representations before registration, where the proposed similarity matrix downsampling method is used to filter out the points with low similarity and retain reliable representations, and thus reduce the side effects of overlap estimation errors on the registration. Besides, compared with previous similarity-based and score-based overlap estimation methods, we use the dual-branch structure to combine the benefits of both, which is less sensitive to noise. We perform overlap estimation experiments and registration experiments on the ModelNet40 dataset, outdoor large scene dataset KITTI, and natural data Stanford Bunny dataset. The experimental results demonstrate that our method is superior to other partial registration methods. Our code is available at https://github.com/superYuezhang/RORNet. Yue Wu 0004, Yue Zhang 0040, Wenping Ma 0001, Maoguo Gong, Xiaolong Fan, Mingyang Zhang 0002, A. K. Qin 0001, Qiguang Miao |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2024 | Semisupervised Change Detection Based on Bihierarchical Feature Aggregation and Extraction NetworkabstractWith the rapid development of remote sensing (RS) technology, high-resolution RS image change detection (CD) has been widely used in many applications. Pixel-based CD techniques are maneuverable and widely used, but vulnerable to noise interference. Object-based CD techniques can effectively utilize the abundant spectrum, texture, shape, and spatial information but easy-to-ignore details of RS images. How to combine the advantages of pixel-based methods and object-based methods remains a challenging problem. Besides, although supervised methods have the capability to learn from data, the true labels representing changed information of RS images are often hard to obtain. To address these issues, this article proposes a novel semisupervised CD framework for high-resolution RS images, which employs small amounts of true labeled data and a lot of unlabeled data to train the CD network. A bihierarchical feature aggregation and extraction network (BFAEN) is designed to achieve the pixelwise together with objectwise feature concatenation feature representation for the comprehensive utilization of the two-level features. In order to alleviate the coarseness and insufficiency of labeled samples, a confident learning algorithm is used to eliminate noisy labels and a novel loss function is designed for training the model using true- and pseudo-labels in a semisupervised fashion. Experimental results on real datasets demonstrate the effectiveness and superiority of the proposed method. Mingyang Zhang 0002, Tianqi Gao, Maoguo Gong, Shengqi Zhu 0001, Yue Wu 0004, Hao Li 0009 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2024 | EGST: Enhanced Geometric Structure Transformer for Point Cloud RegistrationabstractWe explore the effect of geometric structure descriptors on extracting reliable correspondences and obtaining accurate registration for point cloud registration. The point cloud registration task involves the estimation of rigid transformation motion in unorganized point cloud, hence it is crucial to capture the contextual features of the geometric structure in point cloud. Recent coordinates-only methods ignore numerous geometric information in the point cloud which weaken ability to express the global context. We propose Enhanced Geometric Structure Transformer to learn enhanced contextual features of the geometric structure in point cloud and model the structure consistency between point clouds for extracting reliable correspondences, which encodes three explicit enhanced geometric structures and provides significant cues for point cloud registration. More importantly, we report empirical results that Enhanced Geometric Structure Transformer can learn meaningful geometric structure features using none of the following: (i) explicit positional embeddings, (ii) additional feature exchange module such as cross-attention, which can simplify network structure compared with plain Transformer. Extensive experiments on the synthetic dataset and real-world datasets illustrate that our method can achieve competitive results. Yongzhe Yuan, Yue Wu 0004, Xiaolong Fan, Maoguo Gong, Wenping Ma 0001, Qiguang Miao |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2023 | Exploring Dual Representations in Large-Scale Point Clouds: A Simple Weakly Supervised Semantic Segmentation FrameworkabstractExisting work shows that 3D point clouds produce only about a 4% drop in semantic segmentation even at 1% random point annotation, which inspires us to further explore how to achieve better results at lower cost. As scene point clouds provide position and color information and often used in tandem as the only input, with little work going into segmentation by fusing information from dual spaces. To optimize point cloud representations, we propose a novel framework for the dual representation query network (DRQNet). The proposed framework partitions the input point cloud into position and color spaces, using the separately extracted geometric structure and semantic context to create an internal supervisory mechanism that bridges the dual spaces and fuses the information. Adopting sparsely annotated points as the query set, DRQNet provide guidance and perceptual information for multi-stage point clouds through random sampling. More, to differentiate and enhance the features generated by local neighbourhoods within multiple perceptual fields, we design a representation selection module to identify the contributions made by the position and color of each query point, and weight them adaptively according to reliability. The proposed DRQNet is robust to point cloud analysis and eliminates the effects of irregularities and disorder. Our method achieves significant performance gains on three mainstream benchmarks. Yue Wu 0004, Maoguo Gong, Qiguang Miao, Wenping Ma 0001 |
ACM Multimedia | 2 |
| 2023 | Unbalanced Multi-view Deep LearningabstractMost existing multi-view learning methods assume that the dimensions of different views are similar. In real-world applications, it is often the case that the dimension of a view may be extremely small compared with these of other views, resulting in an unbalanced multi-view learning problem. Previous methods for this problem have at least one of the following drawbacks: (1) despising the information of low dimensional views; (2) constructing balanced view-specific inter-instance similarity graphs or employing decision-level fusion, which cannot well learn multi-level inter-view correlations and is limited to category-related tasks such as clustering. To eliminate all these drawbacks, we present an Unbalanced Multi-view Deep Learning (UMDL) method. Considering a low dimensional view usually contains multiple patterns, we construct an overcomplete dictionary with its atoms exceeding the dimension of the original data. We transfer the original data into a combination of atoms and obtain a higher dimensional representation. We propose a sparse multi-view fusion paradigm to explicitly capture the complementarity of multi-view data in a flexible manner. Moreover, we construct positive and negative examples via balanced similarity graphs and employ contrastive learning to train UMDL in a self-supervised manner. Experiments conducted on a toy example and 7 balanced/unbalanced datasets show that UMDL outperforms baseline methods and can be well applied to downstream classification and segmentation tasks. The code is released at https://github.com/xdmvteam/UMDL. Ziyu Guan, Wei Zhao 0019, Yue Wu 0004, Jianxin Li 0001 |
ACM Multimedia | 6 |
| 2023 | Heterogeneous Multi-Party Learning With Data-Driven Network SamplingabstractMulti-party learning provides an effective approach for training a machine learning model, e.g., deep neural networks (DNNs), over decentralized data by leveraging multiple decentralized computing devices, subjected to legal and practical constraints. Different parties, so-called local participants, usually provide heterogenous data in a decentralized mode, leading to non-IID data distributions across different local participants which pose a notorious challenge for multi-party learning. To address this challenge, we propose a novel heterogeneous differentiable sampling (HDS) framework. Inspired by the dropout strategy in DNNs, a data-driven network sampling strategy is devised in the HDS framework, with differentiable sampling rates which allow each local participant to extract from a common global model the optimal local model that best adapts to its own data properties so that the size of the local model can be significantly reduced to enable more efficient inference. Meanwhile, co-adaptation of the global model via learning such local models allows for achieving better learning performance under non-IID data distributions and speeds up the convergence of the global model. Experiments have demonstrated the superiority of the proposed method over several popular multi-party learning techniques in the multi-party settings with non-IID data distributions. Maoguo Gong, Yuan Gao 0019, Yue Wu 0004, Yuanqiao Zhang, A. K. Qin 0001, Yew-Soon Ong |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2023 | Markov clustering regularized multi-hop graph neural network
Xiaolong Fan, Maoguo Gong, Yue Wu 0004 |
Pattern Recognit. | 3 |
| 2023 | Features kept generative adversarial network data augmentation strategy for hyperspectral image classification
Mingyang Zhang 0002, Zhaoyang Wang 0003, Maoguo Gong, Yue Wu 0004, Hao Li 0009 |
Pattern Recognit. | 5 |
| 2023 | Self-structured pyramid network with parallel spatial-channel attention for change detection in VHR remote sensed imagery
Mingyang Zhang 0002, Hanhong Zheng, Maoguo Gong, Yue Wu 0004, Hao Li 0009, Xiangming Jiang |
Pattern Recognit. | 4 |
| 2023 | SACF-Net: Skip-Attention Based Correspondence Filtering Network for Point Cloud RegistrationabstractRigid registration is a transformation estimation problem between two point clouds. The two point clouds captured may partially overlap owing to different viewpoints and acquisition times. Some previous correspondence matching based methods utilize an encoder-decoder network to carry out partial-to-partial registration task and adopt a skip-connection structure to convey information between the encoder and decoder. However, equally revisiting them with skip-connection may introduce the information redundancy, and limit the feature learning ability of the entire network. To address these problems, we propose a skip-attention based correspondence filtering network (SACF-Net) for point cloud registration. A novel feature interaction mechanism is designed to utilize both low-level geometric information and high-level context-aware information to enhance the original pointwise matching map. Additionally, a skip-attention based correspondence filtering method is proposed to selectively revisits features in the encoder at different resolutions, allowing the decoder to extract high-quality correspondences within overlapping regions. We conduct comprehensive experiments on indoor and outdoor scene datasets, and the results show that the proposed SACF-Net yields unprecedented performance improvements. Yue Wu 0004, Xidao Hu, Yue Zhang 0040, Maoguo Gong, Wenping Ma 0001, Qiguang Miao |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2023 | INENet: Inliers Estimation Network With Similarity Learning for Partial Overlapping RegistrationabstractPoint cloud registration is a key problem in the application of computer vision to robotics, autopilot and other fields. However, because the object is partially covered up or the resolution of 3D scanners is different, point clouds collected by the same sense may be inconsistent and even incomplete. Inspired by the recently proposed learning-based approaches, we propose Inliers Estimation Network (INENet) which includes a self-designed threshold prediction network and a probability estimation network with adaptive similarity mutual attention to help to find the overlapping area of the point clouds. In order to solve the above problems, we divide the partially overlapping point cloud registration task into two sub-tasks: overlapping areas detection and registration. The threshold prediction network can automatically calculate the threshold according to the input point clouds, and then the probability estimation network estimates the overlapping points by using threshold. The advantages of the proposed approach include: (1) threshold prediction network avoids bias and the complexity of manually adjusting the threshold. (2) Probability estimation network with similarity matrix can deeply fuse the information between a pair of point clouds, which is helpful to improve the accuracy. (3) INENet can be easily integrated into other overlapping region sensitive algorithms and without adjusting parameters. We conduct experiments on the ModelNet40, S3DIS and 3DMatch data sets. Specifically, the rotation error of the registration algorithm integrated with INENet is improved by at least 25% compared with direct partial overlap registration, our method improves the$F_{1} $score by 5% and has better anti-noise ability compared with the existing overlap detection methods, showing the effectiveness of the proposed method. Yue Wu 0004, Yue Zhang 0040, Xiaolong Fan, Maoguo Gong, Qiguang Miao, Wenping Ma 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2023 | Multiform Ensemble Self-Supervised Learning for Few-Shot Remote Sensing Scene ClassificationabstractSelf-supervised learning is an effective way to solve model collapse for few-shot remote sensing scene classification (FSRSSC). However, most self-supervised contrastive learning auxiliary tasks perform poorly on the high interclass similarity problem in FSRSSC. Furthermore, it is time-consuming and computationally expensive to obtain the best combination among numerous self-supervised auxiliary tasks. In practical applications, we may encounter difficulties in remote sensing data acquisition and labeling, while most FSRSSC studies only focus on the former. To alleviate the above problems, we propose a multiform ensemble self-supervised learning (MES2L) framework for FSRSSC in this article. Based on the transfer learning-based few-shot scheme, we design a novel global–local contrastive learning auxiliary task to solve the low interclass separability problem. The self-attention mechanism is designed in the local contrast features to investigate the intrinsic associations between different remote sensing scene objectives. We also present a multiform ensemble enhancement (MEE) training method. Ensemble enhancement involves the concatenation of features extracted from different backbones trained by a combination of multiform self-supervised auxiliary tasks. MEE can not only be regarded as a more straightforward alternative to knowledge distillation but also can achieve an effective compromise between expensive computational cost and classification accuracy. In addition, we provide two scene classification schemes of inductive and transductive settings, corresponding to solving the difficulties of remote sensing data acquisition and labeling. The proposed network achieves state-of-the-art results on three benchmark FSRSSC datasets. The potential of the MES2L framework is also demonstrated in combination with classical metalearning-based and metric learning-based few-shot algorithms. Jianzhao Li, Maoguo Gong, Huilin Liu, Yourun Zhang, Mingyang Zhang 0002, Yue Wu 0004 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2023 | Cross-Domain Self-Taught Network for Few-Shot Hyperspectral Image ClassificationabstractIn recent years, deep learning models, which possess powerful feature extraction abilities, have achieved remarkable success in the classification of hyperspectral images (HSIs). Nevertheless, a common challenge faced by most deep learning models, including few-shot learning models, is the scarcity of valid labeled samples. To address this issue, we propose a cross-domain self-taught network (CDSTN) for few-shot hyperspectral image classification. The proposed CDSTN merges domain adaptation and semi-supervised self-taught strategy to implement the few-shot learning, which utilizes adequate labeled and unlabeled samples from source as well as target domain respectively. For the feature information extraction of HSI, we propose a deep spatial-spectral feature embedded extractor composed of four residual blocks and a channel attention module. Additionally, a set of domain classifiers are introduced behind each residual block for the purpose of domain alignment by extracting more domain information at different depths of the network. Finally, plenty of unlabeled samples are assigned with pseudo labels through the trained network, and a pseudo label refinement module is designed to select the most confident pseudo label sample for each class to further enrich the labeled database of target domain. Experiments conducted on four widely used benchmark HSI data sets demonstrate that CDSTN can obtain superior and stable performance with limited labeled samples compared with some state of the arts. Mingyang Zhang 0002, Hao Liu 0123, Maoguo Gong, Hao Li 0009, Yue Wu 0004, Xiangming Jiang |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2023 | Maximizing Mutual Information Across Feature and Topology Views for Representing GraphsabstractRecently, maximizing mutual information has emerged as a powerful tool for unsupervised graph representation learning. Existing methods are typically effective in capturing graph information from the topology view but consistently ignore the node feature view. To circumvent this problem, we propose a novel method by exploiting mutual information maximization across feature and topology views. Specifically, we first construct the feature graph to capture the underlying structure of nodes in feature spaces by measuring the distance between pairs of nodes. Then we use a cross-view representation learning module to capture both local and global information content across feature and topology views on graphs. To model the information shared by the feature and topology spaces, we develop a common representation learning module by using mutual information maximization and reconstruction loss minimization. Here, minimizing reconstruction loss forces the model to learn the shared information of feature and topology spaces. To explicitly encourage diversity between graph representations from the same view, we also introduce a disagreement regularization to enlarge the distance between representations from the same view. Experiments on synthetic and real-world datasets demonstrate the effectiveness of integrating feature and topology views. In particular, compared with the previous supervised methods, the proposed method achieves comparable or even better performance under the unsupervised representation and linear evaluation protocol. Xiaolong Fan, Maoguo Gong, Yue Wu 0004, Hao Li 0009 |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2023 | Propagation Enhanced Neural Message Passing for Graph Representation LearningabstractGraph Neural Network (GNN) is capable of applying deep neural networks to graph domains. Recently, Message Passing Neural Networks (MPNNs) have been proposed to generalize several existing graph neural networks into a unified framework. For graph representation learning, MPNNs first generate discriminative node representations using the message passing function and then read from the node representation space to generate a graph representation using the readout function. In this paper, we analyze the representation capacity of the MPNNs for aggregating graph information and observe that the existing approaches ignore the self-loop for graph representation learning, leading to limited representation capacity. To alleviate this issue, we introduce a simple yet effective propagation enhanced extension, Self-Connected Neural Message Passing (SC-NMP), which aggregates the node representations of the current step and the graph representation of the previous step. To further improve the information flow, we also propose a Densely Self-Connected Neural Message Passing (DSC-NMP) that connects each layer to every other layer in a feed-forward fashion. Both proposed architectures are applied at each layer and the graph representation can then be used as input into all subsequent layers. Remarkably, combining these two architectures with existing GNN variants can improve these models’ performance for graph representation learning. Extensive experiments on various benchmark datasets strongly demonstrate the effectiveness, leading to superior performance for graph classification and regression tasks. Xiaolong Fan, Maoguo Gong, Yue Wu 0004, A. K. Qin 0001, Yu Xie 0009 |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2022 | Ternary Change Detection in SAR Images Based on Bi-hierarchical SDAE and Bayesian OptimizationabstractIn this paper, we propose a new change detection method of multi-temporal synthetic aperture radar (SAR) images. Due to the ability of extracting key feature of images and robustness to noise, stacked denoising auto encoder (SDAE) has been widely used in remote sensing. However, the single SDAE stills has some limitations to handle with the speckle noise of SAR images. Therefore, we propose a new structure Bi-hierarchical SDAE for feature extraction. The first level of SDAE denoises the original image and reconstructs the difference map, and the second level extracts the superpixel-based difference features for classification. Besides, Bayesian optimization effectively improves the classification performance of feature classifier. The experimental results of the datasets in this paper show that the Bi-hierarchical SDAE and Bayesian optimization framework has high accuracy and proves its effectiveness. Zhuping Hu, Tianqi Gao, Hao Li 0009, Maoguo Gong, Yue Wu 0004, Jieyi Liu, Jiao Shi |
IJCNN | 5 |
| 2022 | Evolutionary Multitasking CNN Architecture Search for Hyperspectral Image ClassificationabstractIn recent years, convolutional neural networks (CNNs) have shown excellent effectiveness on hyperspectral image classification (HSI) tasks. However, it is a challenge to design a suitable CNN architecture to obtain great performance according to different tasks. Different from the traditional manual design, in this paper, an evolutionary multitasking CNN architecture search framework for HSI classification is proposed to search the optimal architectures and accomplish classification of different tasks simultaneously. Through encoding the CNN architectures, the proposed algorithm is able to achieve global search in the same search space and select well-adapted individuals for evolution. In the evolutionary multitasking environment, information can be transferred between and within tasks, which can accelerate the convergence and explore good architectures through beneficial transfer. In the experiments, the effectiveness of the proposed method is demonstrated by the comparison with different methods on two common data sets. Yiting Liu 0004, Hao Li 0009, Maoguo Gong, Jieyi Liu, Yue Wu 0004, Mingyang Zhang 0002, Jiao Shi |
IJCNN | 5 |
| 2022 | An Automatically Layer-Wise Searching Strategy for Channel Pruning Based on Task-Driven Sparsity OptimizationabstractDeep convolutional neural networks (CNNs) have achieved tremendous successes but tend to suffer from high computation costs mainly due to heavy over-parameterization, resulting in the difficulty of directly applying them to the ever-growing application demands based on low-end edge devices with strong power restriction and real-time inference requirement. Recently, there has much research attention devoted to compressing the network via pruning to address this issue. Most of the existing methods rely on some hand-designed pruning rules, which suffer from several limitations. Firstly, manually designed rules are only applicable to limited application scenarios, which can hardly generalize well in a broader scope. And these rules are typically designed based on human experience and via trial and error, and thus highly subjective. Then, channels of different layers in a network may have diverse distributions, which means the same pruning rule is not appropriate for each layer. To address these limitations, we propose a novel channel pruning scheme, in which the task-irrelevant channels are removed in a task-driven manner. Specifically, an adaptively differentiable search module is proposed to find the best pruning rule automatically for different layers in CNNs under sparsity constraints. Besides, we employed knowledge distillation to alleviate the excessive performance loss. Once the training process is finished, a compact network will be obtained by removing channels based on layer-wise pruning rules. We have evaluated the proposed method on some well-known benchmark datasets including CIFAR, MNIST, and ImageNet in comparison to several state-of-the-art pruning methods. Experimental results demonstrate the superiority of our method over the compared ones in terms of both parameters and FLOPs reduction. Kaiyuan Feng, Xia Fei, Maoguo Gong, A. K. Qin 0001, Hao Li 0009, Yue Wu 0004 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2022 | A Multifactorial Optimization Framework Based on Adaptive Intertask Coordinate SystemabstractThe searching ability of the population-based search algorithms strongly relies on the coordinate system on which they are implemented. However, the widely used coordinate systems in the existing multifactorial optimization (MFO) algorithms are still fixed and might not be suitable for various function landscapes with differential modalities, rotations, and dimensions; thus, the intertask knowledge transfer might not be efficient. Therefore, this article proposes a novel intertask knowledge transfer strategy for MFOs implemented upon an active coordinate system that is established on a common subspace of two search spaces. The proper coordinate system might identify some common modality in a proper subspace to some extent. In this article, to seek the intermediate subspace, we innovatively introduce the geodesic flow that starts from a subspace, reaching another subspace in unit time. A low-dimension intermediate subspace is drawn from a uniform distribution defined on the geodesic flow, and the corresponding coordinate system is given. The intertask trial generation method is applied to the individuals by first projecting them on the low-dimension subspace, which reveals the important invariant features of the multiple function landscapes. Since intermediate subspace is generated from the major eigenvectors of tasks' spaces, this model turns out to be intrinsically regularized by neglecting the minor and small eigenvalues. Therefore, the transfer strategy can alleviate the influence of noise led by redundant dimensions. The proposed method exhibits promising performance in the experiments. Zedong Tang, Maoguo Gong, Yue Wu 0004, A. K. Qin 0001, Kay Chen Tan |
IEEE Trans. Cybern. | 3 |
| 2022 | Two-Path Aggregation Attention Network With Quad-Patch Data Augmentation for Few-Shot Scene ClassificationabstractThe few-shot scene classification is dedicated to identifying unseen remote sensing classes when only a very small number of labeled samples are available for reference. Most of the existing few-shot scene classification methods are based on meta-learning and employ the episodic learning for training, which lacks the consideration for the utilization of data efficiency. In this paper, instead of designing sophisticated meta-learning based algorithms, we are committed to training a feature extractor with good generalization performance and strong feature extraction capability. Specifically, we propose a novel two-path aggregation attention network with quad-patch data augmentation, called DANet, to solve the problem of few-shot scene classification from both data and architecture aspects. In terms of data, we design a new data augmentation strategy named quad-patch augmentation. We utilize the characteristics of remote sensing images to chunk and reassemble any existing data, thereby generating pseudo-new data to enrich the training set. In terms of architecture, we present a two-path aggregation attention module that makes it easier for the model to focus on the key clues in a targeted manner. The comparative experiments in natural image datasets and remote sensing image datasets demonstrate the effectiveness of our two innovations. In addition, DANet achieves competitive or state-of-the-art (SOTA) results on three benchmark scene classification datasets. Maoguo Gong, Jianzhao Li, Yourun Zhang, Yue Wu 0004, Mingyang Zhang 0002 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | Deep Neural Message Passing With Hierarchical Layer Aggregation and Neighbor NormalizationabstractAs a unified framework for graph neural networks, message passing-based neural network (MPNN) has attracted a lot of research interest and has been shown successfully in a number of domains in recent years. However, because of over-smoothing and vanishing gradients, deep MPNNs are still difficult to train. To alleviate these issues, we first introduce a deep hierarchical layer aggregation (DHLA) strategy, which utilizes a block-based layer aggregation to aggregate representations from different layers and transfers the output of the previous block to the subsequent block, so that deeper MPNNs can be easily trained. Additionally, to stabilize the training process, we also develop a novel normalization strategy, neighbor normalization (NeighborNorm), which normalizes the neighbor of each node to further address the training issue in deep MPNNs. Our analysis reveals that NeighborNorm can smooth the gradient of the loss function, i.e., adding NeighborNorm makes the optimization landscape much easier to navigate. Experimental results on two typical graph pattern-recognition tasks, including node classification and graph classification, demonstrate the necessity and effectiveness of the proposed strategies for graph message-passing neural networks. Xiaolong Fan, Maoguo Gong, Zedong Tang, Yue Wu 0004 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2022 | Commonality Autoencoder: Learning Common Features for Change Detection From Heterogeneous ImagesabstractChange detection based on heterogeneous images, such as optical images and synthetic aperture radar images, is a challenging problem because of their huge appearance differences. To combat this problem, we propose an unsupervised change detection method that contains only a convolutional autoencoder (CAE) for feature extraction and the commonality autoencoder for commonalities exploration. The CAE can eliminate a large part of redundancies in two heterogeneous images and obtain more consistent feature representations. The proposed commonality autoencoder has the ability to discover common features of ground objects between two heterogeneous images by transforming one heterogeneous image representation into another. The unchanged regions with the same ground objects share much more common features than the changed regions. Therefore, the number of common features can indicate changed regions and unchanged regions, and then a difference map can be calculated. At last, the change detection result is generated by applying a segmentation algorithm to the difference map. In our method, the network parameters of the commonality autoencoder are learned by the relevance of unchanged regions instead of the labels. Our experimental results on five real data sets demonstrate the promising performance of the proposed framework compared with several existing approaches. Yue Wu 0004, Yongzhe Yuan, A. K. Qin 0001, Qiguang Miao, Maoguo Gong |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2021 | SKFAC: Training Neural Networks With Faster Kronecker-Factored Approximate CurvatureabstractThe bottleneck of computation burden limits the widespread use of the 2nd order optimization algorithms for training deep neural networks. In this paper, we present a computationally efficient approximation for natural gradient descent, named Swift Kronecker-Factored Approximate Curvature (SKFAC), which combines Kronecker factorization and a fast low-rank matrix inversion technique. Our research aims at both fully connected and convolutional layers. For the fully connected layers, by utilizing the low-rank property of Kronecker factors of Fisher information matrix, our method only requires inverting a small matrix to approximate the curvature with desirable accuracy. For convolutional layers, we propose a way with two strategies to save computational efforts without affecting the empirical performance by reducing across the spatial dimension or receptive fields of feature maps. Specifically, we propose two effective dimension reduction methods for this purpose: Spatial Subsampling and Reduce Sum. Experimental results of training several deep neural networks on Cifar-10 and ImageNet-1k datasets demonstrate that SKFAC can capture the main curvature and yield comparative performance to K-FAC. The proposed method bridges the wall-clock time gap between the 1st and 2nd order algorithms. Zedong Tang, Fenlong Jiang, Maoguo Gong, Hao Li 0009, Yue Wu 0004, Fan Yu 0004, Zidong Wang 0010, Min Wang 0037 |
CVPR | 5 |
| 2021 | Regularized Evolutionary Multitask Optimization: Learning to Intertask Transfer in Aligned SubspaceabstractThis article proposes a novel and computationally efficient explicit intertask information transfer strategy between optimization tasks by aligning the subspaces. In evolutionary multitasking, the tasks might have biases embedded in function landscapes and decision spaces, which often causes the threat of predominantly negative transfer. However, the complementary information among different tasks can give an enhanced performance of solving complicated problems when properly harnessed. In this article, we distill this insight by introducing an intertask knowledge transfer strategy implemented in the low-dimension subspaces via a learnable alignment matrix. Specifically, to unveil the significant features of the function landscapes, the task-specific low-dimension subspaces is established based on the distribution information of subpopulations possessed by tasks, respectively. Next, the alignment matrix between pairwise subspaces is learned by minimizing the discrepancies of the subspaces. Given the aligned subspaces by applying the alignment matrix to subspaces' base vectors, the individuals from different tasks are then projected into aligned subspaces and reproduce therein. Moreover, since this method only considers the leading eigenvectors, it turns out to be intrinsically regularized and noise-insensitive. Comprehensive experiments are conducted on the synthetic and practical benchmark problems so as to assess the efficacy of the proposed method. According to the experimental results, the proposed method exhibits a superior performance compared with existing evolutionary multitask optimization algorithms. Zedong Tang, Maoguo Gong, Yue Wu 0004, Yu Xie 0009 |
IEEE Trans. Evol. Comput. | 3 |
| 2020 | Endmember Selection of Hyperspectral Images based on Evolutionary MultitaskabstractEndmember selection of hyperspectral images is a practical yet difficult task due to the high spectral resolution and low spatial resolution of the hyperspectral cameras. The paradigm of multitask optimization has been investigated over two decades, which aim to handle multiple tasks simultaneously. To address these issues, we propose a novel multitasking framework based on multiobjective optimization evolutionary algorithm based on decomposition (MOEA/D). Specifically, we use a single population to simultaneously perform multiple subset selection tasks and apply it to a specific scene-the endmember selection of hyperspectral images. It is natural to consider that pixels in a homogeneous region of hyperspectral image as a task. Then, a within-task and between-task genetic transfer operator is constructed to reinforce the exchange of genetic material belonging to the same or different tasks for better and quicker search of the decision space. After that, this algorithm obtains a set of nondominated solutions for better decision of the active endmembers. Experiments on hyperspectral datasets show the effectiveness of our method in finding the real active endmembers. Hao Li 0009, Yue Wu 0004, Shanfeng Wang, Maoguo Gong |
CEC | 3 |
| 2020 | Gated Graph Pooling with Self-Loop for Graph ClassificationabstractGraph classification is a practical problem in many different domains including bioinformatics, chemoinformatics, social network analysis, and etc. For the graph classification task, the existing graph neural network approaches usually generate graph features using graph pooling at each step. However, this strategy of pooling only at the current step ignores the impact of self-loop. To eliminate this limitation, we propose a novel self-loop graph pooling strategy that can utilize the node information of the current step and the graph representation information of the previous step to generate an effective representation for the graph classification task. Further to measure the importance of self-loop, we also develop a gated approach, gated graph pooling with self-loop, that utilizes the simple fusion gate to enhance the representation capacity of the model. We evaluate our model on common benchmark datasets and experimental results have demonstrated the superior performance improvement on predictive accuracy. Xiaolong Fan, Maoguo Gong, Hao Li 0009, Yue Wu 0004, Shanfeng Wang |
IJCNN | 4 |
| 2020 | Multi-objective optimization for location-based and preferences-aware recommendation
Shanfeng Wang, Maoguo Gong, Yue Wu 0004, Mingyang Zhang 0002 |
Inf. Sci. | 3 |
| 2019 | Multipopulation Optimization for Multitask OptimizationabstractCurrently, the most of multitask evolutionary algorithms views multiple tasks as factors influencing the evolution of individuals. However, this consideration causes difficulty to assign fitness to individuals, because an individual which performs well on one task can have a bad performance on another task. To avoid this difficulty, this paper proposes a novel multipopulation technique for multitask optimization (MPMTO). The novelty of MPMTO is that it can solve the multiple tasks via a simple and straightforward method by corresponding each population to a task. By this way, the fitness assignment issue can be addressed by just assigning the objective value of the corresponding task to individuals. MPMTO is a general technique so that existing population-based optimization algorithms can be used in each population. This paper uses differential evolutionary algorithm in each population and develops a multipopulation multitask differential evolutionary optimization (mMTDE) based on the proposed multipopulation technique. mMTDE features that each population can use the other populations as the additional knowledge source to create an overlapping population, allowing the populations share information. By this way, the population can improve the efficacy and accuracy of solving multiple tasks. Moreover, the successful inter-task offspring can immigrate back to the corresponding population to fully utilize the inter-task knowledge. We have compared the proposed method with other state-of-the-art methods on benchmark multitask problems. The experimental results show the superiority of the proposed method which could utilizes efficiently the searching knowledge of multiple tasks. Zedong Tang, Maoguo Gong, Fenlong Jiang, Hao Li 0009, Yue Wu 0004 |
CEC | 5 |
| 2019 | High-order graph matching based on ant colony optimization
Yue Wu 0004, Maoguo Gong, Wenping Ma 0001, Shanfeng Wang |
Neurocomputing | 1 |
| 2019 | A two-step personalized location recommendation based on multi-objective immune algorithm
Bingrui Geng, Licheng Jiao, Maoguo Gong, Lingling Li 0002, Yue Wu 0004 |
Inf. Sci. | 5 |
| 2019 | Multimodal Remote Sensing Image Registration Based on Image Transfer and Local FeaturesabstractAutomatic registration is still a challenging problem for multimodal remote sensing images including optical, light detection and ranging, synthetic aperture radar images, and so on. Due to the differences in imaging principles, the gray value, texture, and landscape characteristic of these images are different in the local area. This also makes it difficult to obtain satisfactory results for the conventional image registration methods. In order to achieve registration of multimodal images to obtain complementary information, we apply the transfer algorithm based on a deep image analogy to the preprocessing of image registration. It eliminates the differences in multimodal remote sensing images by blending the original image structure and texture. The conventional local feature-based method is applied to match the original and generated images. Correspondences are increased and the registration error is reduced. The experiments demonstrate that our method can effectively deal with multimodal data and produce more accurate results. The algorithm is based on the joint of image deep semantic features and indirectly achieves matching of the original image pair. It provides a new solution to the problem of multimodal images registration. Jun Zhang 0045, Wenping Ma 0001, Yue Wu 0004, Licheng Jiao |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2019 | A Novel Two-Step Registration Method for Remote Sensing Images Based on Deep and Local FeaturesabstractAutomatic remote sensing image registration has achieved great accomplishment. However, it is still a vital challenging problem to develop a robust and accurate registration method due to the negative effects of noise and imaging differences between images. For these images, it is difficult to guarantee the accuracy and robustness at the same time for one-step registration methods. To address this issue, we introduce an effective coarse-to-fine strategy and develop a new two-step registration method based on deep and local features in this paper. The first step is to calculate the approximate spatial relationship, which is obtained by a convolutional neural network. This step makes full use of the deep features to match and can generate stable results. For the second step, a matching strategy considering spatial relationship is applied to the local feature-based method. In addition, this step adopts more accurate features in location to adjust the results of the previous step. A variety of homologous and multimodal remote sensing images, including optical, synthetic aperture radar, and general map images, are used to evaluate the proposed method. The comparison experiments demonstrate that our method can apparently increase the correct correspondences, can improve the ratio of correct correspondences, and is highly robust and accurate. Wenping Ma 0001, Jun Zhang 0045, Yue Wu 0004, Licheng Jiao, Hao Zhu 0009, Wei Zhao 0014 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2019 | Unsupervised Feature Extraction in Hyperspectral Images Based on Wasserstein Generative Adversarial NetworkabstractFeature extraction (FE) is a crucial research area in hyperspectral image (HSI) processing. Recently, due to the powerful ability of deep learning (DL) to extract spatial and spectral features, DL-based FE methods have shown great potentials for HSI processing. However, most of the DL-based FE methods are supervised, and the training of them suffers from the absence of labeled samples in HSIs severely. The training issue of supervised DL-based FE methods limits their application on HSI processing. To address this issue, in this paper, a novel modified generative adversarial network (GAN) is proposed to train a DL-based feature extractor without supervision. The designed GAN consists of two components, which are a generator and a discriminator. The generator can focus on the learning of real probability distributions of data sets and the discriminator can extract spatial-spectral features with superior invariance effectively. In order to learn upsampling and downsampling strategies adaptively during FE, the proposed generator and discriminator are designed based on a fully deconvolutional subnetwork and a fully convolutional subnetwork, respectively. Moreover, a novel min-max cost function is designed for training the proposed GAN in an end-to-end fashion without supervision, by utilizing the zero-sum game relationship between the generator and discriminator. Besides, the proposed modified GAN replaces the original Jensen-Shannon divergence with the Wasserstein distance, aiming to mitigate the unstability and difficulty of the training of GAN frameworks. Experimental results on three real data sets validate the effectiveness of the proposed method. Mingyang Zhang 0002, Maoguo Gong, Yishun Mao, Jun Li 0009, Yue Wu 0004 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2018 | Feature Hashing for Network Representation LearningabstractThe goal of network representation learning is to embed nodes so as to encode the proximity structures of a graph into a continuous low-dimensional feature space. In this paper, we propose a novel algorithm called node2hash based on feature hashing for generating node embeddings. This approach follows the encoder-decoder framework. There are two main mapping functions in this framework. The first is an encoder to map each node into high-dimensional vectors. The second is a decoder to hash these vectors into a lower dimensional feature space. More specifically, we firstly derive a proximity measurement called expected distance as target which combines position distribution and co-occurrence statistics of nodes over random walks so as to build a proximity matrix, then introduce a set of T different hash functions into feature hashing to generate uniformly distributed vector representations of nodes from the proximity matrix. Compared with the existing state-of-the-art network representation learning approaches, node2hash shows a competitive performance on multi-class node classification and link prediction tasks on three real-world networks from various domains. Qixiang Wang, Shanfeng Wang, Maoguo Gong, Yue Wu 0004 |
IJCAI | 4 |
| 2018 | PSOSAC: Particle Swarm Optimization Sample Consensus Algorithm for Remote Sensing Image RegistrationabstractImage registration is an important preprocessing step for many remote sensing image processing applications, and its result will affect the performance of the follow-up procedures. Establishing reliable matches is a key issue in point matching-based image registration. Due to the significant intensity mapping difference between remote sensing images, it may be difficult to find enough correct matches from the tentative matches. In this letter, particle swarm optimization (PSO) sample consensus algorithm is proposed for remote sensing image registration. Different from random sample consensus (RANSAC) algorithm, the proposed method directly samples the modal transformation parameter rather than randomly selecting tentative matches. Thus, the proposed method is less sensitive to the correct rate than RANSAC, and it has the ability to handle lower correct rate and more matches. Meanwhile, PSO is utilized to optimize parameter as its efficiency. The proposed method is tested on several multisensor remote sensing image pairs. The experimental results indicate that the proposed method yields a better registration performance in terms of both the number of correct matches and aligning accuracy. Yue Wu 0004, Qiguang Miao, Wenping Ma 0001, Maoguo Gong, Shanfeng Wang |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2017 | Memetic algorithm based location and topic aware recommender system
Shanfeng Wang, Maoguo Gong, Haoliang Li, Yue Wu 0004 |
Knowl. Based Syst. | 5 |
| 2017 | Remote Sensing Image Registration With Modified SIFT and Enhanced Feature MatchingabstractThe scale-invariant feature transform algorithm and its many variants are widely used in feature-based remote sensing image registration. However, it may be difficult to find enough correct correspondences for remote image pairs in some cases that exhibit a significant difference in intensity mapping. In this letter, a new gradient definition is introduced to overcome the difference of image intensity between the remote image pairs. Then, an enhanced feature matching method by combining the position, scale, and orientation of each keypoint is introduced to increase the number of correct correspondences. The proposed algorithm is tested on multispectral and multisensor remote sensing images. The experimental results show that the proposed method improves the matching performance compared with several state-of-the-art methods in terms of the number of correct correspondences and aligning accuracy. Wenping Ma 0001, Zelian Wen, Yue Wu 0004, Licheng Jiao, Maoguo Gong, Yafei Zheng, Liang Liu 0001 |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2017 | Remote Sensing Image Registration Based on Multifeature and Region DivisionabstractAlthough many feature-based registration methods have been proposed, automatic image registration is still a challenging task due to the influence of various conditions and uncertain difficulties for remote sensing images. In this letter, a novel image registration method, including two types of feature detectors and a region boundary constraint strategy for matching, is proposed. Two types of features detected by scale-invariant feature transform and Harris operators have advantages of keeping different structural information in the image and increasing the number of keypoints for later matching. Afterward, a region boundary constraint strategy based on the image sketch map is utilized in matching step. This strategy restricts the detected two types of features in their respective structural region and nonstructural region to reduce the incorrect correspondences. Experimental results demonstrate the superiority of our proposed registration algorithm compared with other research works in terms of correct matching number and aligning accuracy. Wenping Ma 0001, Yue Wu 0004, Yafei Zheng, Zelian Wen, Liang Liu 0001 |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2016 | Multi-objective population-based incremental learning for community structure detectionabstractThe community structure detection of complex networks has become a hot topic in the past several years. In this paper, a new discrete framework of population-based incremental learning for complex networks problem is proposed. Based on the proposed discrete framework, a novel multi-objective population-based incremental learning algorithm is proposed to solve community structure detection problem. The proposed algorithm combines population-based incremental learning with the multi-objective evolutionary algorithm based on decomposition, this makes the evolution get directionality and converge fast. In order to discourage premature convergence, a random perturbation operator is adopted. The proposed algorithm has two contradictory objective functions termed as negative ratio association and ratio cut, respectively. The community structure detection results are a set of tradeoff solutions by simultaneous optimizing these two contradictory objective functions. Each of these solutions corresponds to a network community structure at one hierarchical level. Experiments on both real-world and synthetic networks prove the effectiveness of the proposed algorithm. Wenping Ma 0001, Yue Wu 0004, Jie Yun |
CEC | 2 |
| 2016 | Discrete particle swarm optimization for high-order graph matching
Maoguo Gong, Yue Wu 0004, Wenping Ma 0001, A. K. Qin 0001, Zhenkun Wang 0001, Licheng Jiao |
Inf. Sci. | 2 |
| 2015 | A Novel Point-Matching Algorithm Based on Fast Sample Consensus for Image RegistrationabstractRobustness and accuracy are the two main challenging problems in feature-based remote sensing image registration. In this letter, a novel point-matching algorithm is proposed. An improved random sample consensus (RANSAC) algorithm called fast sample consensus (FSC) is proposed. It divides the data set in RANSAC into two parts: the sample set and the consensus set. Sample set has high correct rate and consensus set has a large number of correct matches. An iterative method is put forward to increase the number of correct correspondences. A set of measures has been used to evaluate the registration result. The performance of the proposed method is validated on the evaluation of these measures and the mosaic images. FSC can get more correct matches than RANSAC in less number of iterations, iterative selection of correct matches algorithm and removal of the imprecise points algorithm effectively increase the accuracy of the result. Extensive experimental studies compared with three state-of-the-art methods prove that the proposed algorithm is robust and accurate. Yue Wu 0004, Wenping Ma 0001, Maoguo Gong, Linzhi Su, Licheng Jiao |
IEEE Geosci. Remote. Sens. Lett. | 1 |