Xia Yuan

dblp:69/2223 · DBLP profile ↗
← Back
44ranked-venue papers
6as first author
37since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 28 · 4 first-author · 23 since 2021Artificial intelligence and machine learning · 17 · 3 first-author · 13 since 2021Systems, architecture and hardware · 4 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Beyond Quadratic: Linear-Time Change Detection with RWKV
abstract
Existing paradigms for remote sensing change detection are caught in a trade-off: CNNs excel at efficiency but lack global context, while Transformers capture long-range dependencies at a prohibitive computational cost. This paper introduces ChangeRWKV, a new architecture that reconciles this conflict. By building upon the Receptance Weighted Key Value (RWKV) framework, our ChangeRWKV uniquely combines the parallelizable training of Transformers with the linear-time inference of RNNs. Our approach core features two key innovations: a hierarchical RWKV encoder that builds multi-resolution feature representation, and a novel Spatial-Temporal Fusion Module (STFM) engineered to resolve spatial misalignments across scales while distilling fine-grained temporal discrepancies. ChangeRWKV not only achieves state-of-the-art performance on the LEVIR-CD benchmark, with an 85.46% IoU and 92.16% F1 score, but does so while drastically reducing parameters and FLOPs compared to previous leading methods. This work demonstrates a new, efficient, and powerful paradigm for operational-scale change detection.
Gensheng Pei, Tao Chen 0012, Xia Yuan, Haofeng Zhang 0001, Xiangbo Shu, Yazhou Yao
AAAI4
2026 ScaleGS: Scalable distributed framework for large-scale 3D Gaussian splatting with edge communication
Yong Kou, Xia Yuan, Dening Luo, Yanci Zhang
Perform. Evaluation3
2025 EF2lane: Enhanced Feature Fusion 2D Lane Detection Network In 3d Point Cloud
abstract
Lane detection is one of the core tasks in autonomous driving. Lane detection relies primarily on front-view camera images or the bird’s eye view projection of LiDAR point cloud, but it faces challenges in terms of robustness in complex scenarios. In this paper, we propose a 2D lane detection model based on enhanced feature fusion, EF2Lane. EF2Lane extracts lane features from bird’s-eye view of 3D point cloud, and combines them with Mamba like linear attention (MLLA) to capture global correlations. The global features are then fused with the local features extracted by a convolutional network. Additionally, EF2Lane adopts row-wise inference to efficiently process a sparse point cloud. Experimental results demonstrate that EF2Lane achieves superior performance on the K-Lane and CampusLane datasets.
Yanrui Zhai, Zihui Jing, Xia Yuan
ICIP3
2025 Bilateral Enhanced Complementary Network for Camouflaged Object Detection
abstract
Camouflaged object detection (COD) is a challenging task aimed at identifying and segmenting camouflaged targets that are difficult to distinguish from complex backgrounds. To address the issues of incomplete detection and missing edges in camouflaged targets, this paper proposes a Bilateral Enhanced Complementary Network (BECNet) for COD. The network adopts a two-branch detection method, which is used for object recognition, edge recognition, and texture supervision respectively, to alleviate the feature ambiguity of features extracted from a single branch. Additionally, we introduce a Semantic Amplification Module (SAM) to further extract multi-scale semantic features. To effectively aggregate the discriminative features generated by both branches, we designed a Semantic-Texture Interaction Module (SIM). Finally, we incorporate an Edge Complementary Dual Attention Module (ECDA) during the decoding process to refine the model using edge information. Extensive experiments demonstrate the effectiveness and robustness of BECNet.
Yejing Guo, Xia Yuan, Chunxia Zhao
ICME3
2025 Multi-Source Feature Fusion and Spatio-Temporal Unet for Precipitation Nowcasting
abstract
Precipitation nowcasting is an extremely critical task in the field of weather forecasting, as it facilitates advancements in meteorological observation. Nevertheless, accurate short-term precipitation forecasting remains a significant challenge at present. Traditional methods have relied on physical equations for predictions, which are often computationally consuming. Current deep learning approaches, using CNNs and RNNs, roughly extract the latent features of spatiotemporal data, but the feature extraction process usually overlooks the dynamic changes occurring between prediction image frames. Furthermore, most methods utilize a single precipitation variable as input for predicting future precipitation, neglecting that precipitation events are triggered by multiple meteorological factors. To tackle this issue, we propose a novel neural network model, Multi-Source Feature Fusion and Spatio-Temporal Unet (MFFST-Unet), which utilizes multi-source feature information to guide precipitation forecasting. Additionally, we introduce the Inter-Frame Difference Regularization(IFDR) Loss, which is combined with MSE Loss to optimize the frame stability of model predictions through adaptive weighting. We conducted training and testing on the SEVIR dataset, achieving high-resolution precipitation nowcasting for a one-hour forecast. Experimental results indicate that our MFFST-Unet model surpasses other deep learning methods, achieving a maximum improvement of 34.72% in the CSI precipitation metric, demonstrating its significant practical implications for weather forecasting applications.
Dufu Liu, Xia Yuan, Xi Wu 0004, Jing Hu 0009
IJCNN4
2025 Focusing on Projection-Stable Patch: Cross-View Localization with Geometric-Semantic Alignment
abstract
This paper presents a novel feature alignment strategy for cross-view geo-localization to bridge the perspective gap between ground and satellite images. Existing methods for cross-view geo-localization often overlook factors such as occlusion and distortion errors caused by viewpoint transformation. These issues lead to reduced accuracy in complex scenes. To address this issue, we propose a framework comprising two novel components: a perspective-driven attention fusion (PDAF) module that aligns ground and satellite features through cross-view semantic correlation, effectively preserving structural consistency during view transformation; and a projection-stable patch-guided pose optimizer (PSPG) that enhances geometric reliability by selectively focusing on projection-stable patch to refine pose estimation. The PDAF module mitigates information loss through attention fusion between ground and bird’s-eye-view (BEV) feature maps representations, while the PSPG refines pose estimation by dynamically suppressing unstable features through geometrically unstable token merging. Comprehensive evaluations on KITTI and Ford Multi-AV datasets demonstrate our method’s superiority in orientation estimation and competitive location accuracy compared to state-of-the-art approaches. Qualitative results further confirm the framework’s robustness in complex localization scenarios. The code is available at https://github.com/RobVisLab-NJUST/CVLGSA
Riyu Qin, Xia Yuan
IROS4
2025 BEVPointNet3D: Fusing Bird's Eye View and Point Cloud Features for Robust 3D Lane Detection
abstract
This paper introduces BEVPointNet3D, an innovative 3D lane detection model that effectively integrates Bird’s Eye View (BEV) and point cloud features. The proposed approach addresses the inherent limitations of conventional methods that predominantly rely on the flat-ground assumption. BEVPointNet3D incorporates a 2D encoder to extract preliminary lane information from BEV images. For 3D feature extraction, the model employs a hierarchical local-to-global processing scheme to capture the geometric characteristics of LiDAR point clouds. A novel cross-attention mechanism is implemented to precisely align and integrate the 2D and 3D feature representations. This architectural design not only improves detection accuracy but also strengthens the adaptability and performance of the model in complex driving scenarios. Comprehensive evaluations on the K-Lane and CampusLane datasets demonstrate the superior performance of BEVPoint-Net3D. Notably, the model exhibits exceptional capability in accurately estimating lane spatial positions on steep inclines, thereby providing reliable support for autonomous driving systems in challenging terrain conditions.
Xia Yuan, Yanrui Zhai, Zihui Jing
IROS1
2025 Adaptive mesh-aligned Gaussian Splatting for monocular human avatar reconstruction
abstract
Virtual human avatars are essential for applications such as gaming, augmented reality, and virtual production. However, existing methods struggle to achieve high fidelity reconstruction from monocular input while keeping hardware costs low. Many approaches rely on the SMPL body prior and apply vertex offsets to represent clothed avatars. Unfortunately, excessive offsets often cause misalignment and blurred contours, particularly around clothing wrinkles, silhouette boundaries, and facial regions. To address these limitations, we propose a dual branch framework for human avatar reconstruction from monocular video. A lightweight Vertex Align Net (VAN) predicts per-vertex normal direction offsets on the SMPL mesh to achieve coarse geometric alignment and guide Gaussian-based human avatar modeling. In parallel, we construct a high resolution facial Gaussian branch based on FLAME estimated parameters, with facial regions localized via pretrained detectors. The facial and body renderings are fused using a semantic mask to enhance facial clarity and ensure globally consistent avatar appearance. Experiments demonstrate that our method surpasses state of the art approaches in modeling animatable human avatars with fine grained fidelity.
Hai Yuan, Xia Yuan, Yanli Liu 0002, Guanyu Xing, Zijun Zhou
Graph. Model.2
2025 Active layered topology mapping driven by road intersection
Xia Yuan, Chunxia Zhao
Knowl. Based Syst.2
2025 Objectness scan for efficient vision Mamba
Kai Zhang 0075, Xia Yuan, Chunxia Zhao
Knowl. Based Syst.2
2024 FOTV-HQS: A Fractional-Order Total Variation Model for LiDAR Super-Resolution with Deep Unfolding Network
Huiying Xi, Xia Yuan, Runze Geng, Yongshun Liang, Chunxia Zhao
ACCV (7)2
2024 Procedural Generation of 3D Scenes for Urban Landscape Based on Remote Sensing Images
abstract
3D Reconstruction is a significant research topic in the fields of Computer Graphics, Remote Sensing, Virtual Reality and so on. Particularly, 3D reconstruction of cities is one of the core technologies in urban infrastructure, ecological protection and transportation management. However, City-level scenes are characterized by large-scale and high complexity. The existing large-scale 3D scene reconstruction technology exists problems such as long time and low accuracy. Therefore, it is particularly important to improve the efficiency and realism of large-scale 3D reconstruction of urban scenes. In this paper, a programmed generation of urban largescale 3D scenes based on remote sensing images is proposed. Firstly, semantic-level instance segmentation information is extracted from the urban landscape objects in remote sensing images, then 2D contour information is extracted according to the segmentation results for 3D reconstruction of contour lines, and finally, a highly realistic 3D scene of large-scale urban landscape is rapidly generated according to landscape design rules and Procedural Content Generation Framework (PCG) technology. The method in this paper combines the accurate spatial information in remote sensing images with PCG rapid generation, which can improve the generation speed of 3D scenes and increase the realism of 3D scenes. The experimental results show the effectiveness of this paper’s method.
Shuqin Yang, Haopu Yuan, Chenggang Song, Wenyi Ge, Xia Yuan
AVSS8
2024 AMQGaussian: Efficient 3D Gaussian Representation with Asymmetric Mixed-precision Quantization
abstract
3D Gaussian Splatting (3DGS) has recently gained increasing attention in novel-view scene synthesis. However, it requires millions of 3D Gaussian spheres to achieve high-quality rendered images, leading to substantial GPU resource demands for training and rendering. This paper has developed an efficient framework to address the challenges faced by 3DGS. 1) We propose an asymmetric mixed-precision quantization strategy to efficiently quantize and dequantize the parameters of 3D Gaussian Spheres (excluding position and opacity) and utilized the Straight Through Estimator method to address the issue of non-backpropagatable gradients post-quantization. This approach significantly reduces GPU memory usage and model storage space. 2) Using statistical methods, we identify optimal gradient thresholds to enhance the densification algorithm of 3D Gaussian spheres, thereby further improving performance. 3) To address the speed bottlenecks in parallel processing of low-precision data using CUDA, we introduce a fast atomic operation method, increasing training speed tenfold. Overall, we validate the effectiveness of our framework across various datasets. It reduces GPU usage by 50% during training, decreases GPU usage threefold during rendering, halves the model storage space and maintains image quality comparable to 3DGS.
Yong Kou, Yanci Zhang, Xia Yuan
ISPA4
2024 Multi-Modality Semantic-Shared Cross-View Ground-to-Aerial Localization
Kai Zhang 0075, Xia Yuan, Shuntong Chen, Chunxia Zhao
MMAsia2
2023 Infrared and Visible Image Fusion by Using Multi-Scale Transformation and Fractional-Order Gradient Information
abstract
The fusion of infrared and visible images is hard due to their different modalities. Different from existing methods using the integer-order gradient, we design an optimization model to fuse infrared and visible images using fractional-order gradient information. In this way, the complementary information of the source images can be better preserved. In order to better highlight the target and retain effective details, we use the results of the optimization model as pre-fusion images to guide the final image fusion. For better highlighting the target, we use MDLatLRR to extract the base layer of the pre-fusion image and use it as the base layer of the fused image. In addition, for getting more effective details, we use the pre-fusion image to calculate the weight map, which is used as the tradeoff parameter of the norm optimization problem to get the fused detail layers. Experimental results show that our method can highlight the target better while maintaining effective details. Compared with the current state-of-the-art image fusion methods, our method shows better fusion performance in both subjective and objective evaluation.
Xia Yuan, Chunxia Zhao
ICASSP3
2023 Semantic Mapping of Incremental 3D Point Clouds Based on Multi-Hop Graph Attention Network
abstract
Mapping and Semantic mapping are the important research areas for the autonomous navigation of mobile robots. However, realising point cloud information extraction in a dynamic environment is still a challenge in semantic mapping. To solve this problem, We propose a method called PMGAT-SM(semantic mapping of 3D point clouds based on multi-hop graph attention network) for achieving semantic mapping in 3D point cloud environments. In PMGAT-SM, we designed a point cloud classifier PMGAT, which extracts semantic information of unordered point clouds by constructing a graph. We combined PMGAT with dynamic growing clustering to achieve instance segmentation in complex environments. Our extensive experiments in the KITTI scene show the effectiveness of the semantic mapping model. Meanwhile, compared with popular semantic information extraction models, the PMGAT performs better in fine-grained feature extraction of point cloud segments.
Shuntong Chen, JiaChen Xu, Xia Yuan, Chunxia Zhao
ICIP3
2023 Finding Camouflaged Object Guided by Contour and Attention
abstract
Camouflaged object detection aims to detect objects closely blended into the background. Inspired by visual mechanism of human being, we define detection process as an end-to-end task consisted of locating and refining. To this end, we propose Contour Supervision and Initial Locating Guidance Network (CSIGNet) to effectively segment camouflaged objects from background. Specifically, our method fully explores the contribution of semantic contour to the binary segmentation task. In addition, attention mechanism is used for our final prediction. Experiments show that it has achieved excellent results on public datasets and the model can accurately segment camouflage objects. Codes will be made available at https://github.com/RecKono/CSIGNet.
Junjie Cui, Fengming Sun, Xia Yuan
ICIP3
2023 Contour-Assisted Long-Range Perceptual Network for Camouflaged Instance Segmentation
abstract
High quality instance segmentation has shown emerging significance in computer vision, especially in complex camouflaged situations. For this reason, we propose a single-stage segmentation model named Contour-assisted Long-range Perceptual Network (CLPNet). By following SOLOv2 model, based on deformable transformer, we propose Aggregate and Optimize multi-layer Transformer for generating refining features. Secondly, through iterative optimization, the full use of long-range context dependencies makes the internal of the instance give strong response, and contour of object tightly surrounds the instance. Experiments on camouflaged object detection dateset show that our method reaches 41.7% AP. Compared with the previous instance segmentation method, it is obviously more effective in small object detection.
Junjie Cui, Fengming Sun, Xia Yuan
ICIP3
2023 Feature Enhancement and Fusion for RGB-T Salient Object Detection
abstract
Cross-modal information fusion plays a vital role in the RGB-T salient object detection. Due to RGB and thermal images come from different domains, the modality difference will lead to the unsatisfactory effect of simple feature fusion. How to explore and integrate useful information is the key to the RGB-T saliency detection methods. In this paper, we introduce an Enhancement and Fusion Network. In detail, we propose a Self-modality Feature Enhancement Module that effectively integrate the feature representation of a single modality through global context information. And we propose a Cross-modality Feature Dynamic Fusion Module to realize the effective fusion of cross-modal features in the way of dynamic weighting. Experiments on public datasets show that the proposed method achieves satisfactory results compared with other state-of-the-art salient object detection approaches.
Fengming Sun, Xia Yuan, Chunxia Zhao
ICIP3
2023 RGB-D Road Segmentation Based on Geometric Prior Information
Xia Yuan, YanChao Cui, Chunxia Zhao
PRCV (1)2
2023 Camouflaged Object Segmentation Based on Fractional Edge Perception
Xia Yuan, Junjie Cui, Shuting Yang
PRCV (12)1
2023 Denseformer: A dense transformer framework for person re-identification
abstract
Abstract Transformer has shown its effectiveness and advantage in many computer vision tasks, for example, image classification and object re‐identification (ReID). However, existing vision transformers are stacked layer by layer, lacking direct information exchange among every layer. Inspired by DenseNet, we propose a dense transformer framework (termed Denseformer) that connects each layer to every other layer through class tokens. We demonstrate that Denseformer can consistently achieve better performance on person ReID tasks across datasets (Market‐1501, DukeMTMC, MSMT17, and Occluded‐Duke), only at a negligible increase of computation. We show that Denseformer has several compelling advantages: it pays more attention to the main parts of human bodies and obtains discriminative global features.
Haoyan Ma, Xiang Li 0041, Xia Yuan, Chunxia Zhao
IET Comput. Vis.3
2023 Selective feature fusion network for salient object detection
abstract
Abstract Fully convolutional neural networks have achieved great success in salient object detection, in which the effective use of multi‐layer features plays a critical role. Based on this advantage, many saliency detectors have emerged in recent years, and most of them designed a series of network structures to integrate the multi‐level features generated by the backbone network. However, information in different layer play different roles in saliency object detection, how to integrate them effectively is still a great challenge. In this article, a selective feature fusion network which consists of a selective feature fusion module (SFM) and an attention‐guide hierarchical feature emphasis module (AEM) is proposed. Most of the previous works mainly integrate multi‐level feature by addition and concatenation, as a difference, SFM adaptively selects the important information from the input features in the fusion, which effectively avoids introducing too much redundant information. Besides, AEM combines spatial attention and channel attention to enhance features simply and effectively by hierarchical iteration, and further improve the accuracy of salient object detection. Experiments on five datasets show that the proposed selective feature fusion method achieve satisfactory results when comparing to other state‐of‐the‐art salient object detection approaches.
Fengming Sun, Xia Yuan, Chunxia Zhao
IET Comput. Vis.2
2023 Rich-scale feature fusion network for salient object detection
abstract
Abstract Fully convolutional neural networks‐based salient object detection has recently achieved great success with its performance benefits from the effective use of multi‐layer features. Based on this, most of the existing saliency detectors designed complex network structures to fuse the multi‐level features generated by the backbone network. However, the variable scale and complex shape of the target are always a great challenge for saliency detection tasks. In this paper, the authors propose a Rich‐scale Feature Fusion Network (RFFNet) for salient object detection. The authors design a rich‐scale feature interactive fusion module to obtain more efficient features from the multi‐scale features. Moreover, the global feature enhance module is used to extract features with better characterization for the final saliency prediction. Extensive experiments performed on five benchmark datasets demonstrate that the proposed method can achieve satisfactory results on different evaluation metrics compared to other state‐of‐the‐art salient object detection approaches.
Fengming Sun, Junjie Cui, Xia Yuan, Chunxia Zhao
IET Image Process.3
2023 Two-phase self-supervised pretraining for object re-identification
Haoyan Ma, Xiang Li 0041, Xia Yuan, Chunxia Zhao
Knowl. Based Syst.3
2023 Discriminative and Geometry-Preserving Adaptive Graph Embedding for dimensionality reduction
Jianping Gou, Xia Yuan, Ya Xue, Lan Du 0002, Shuyin Xia, Zhang Yi 0001
Neural Networks2
2023 Divide-and-conquer model based on wavelet domain for multi-focus image fusion
Zhiliang Wu, Hanyu Xuan, Xia Yuan, Chunxia Zhao
Signal Process. Image Commun.4
2023 Intra- and Inter-Class Induced Discriminative Deep Dictionary Learning for Visual Recognition
abstract
Deep dictionary learning (DDL) aims to learn dictionaries at different levels and the deepest level representations. However, existing DDL algorithms impose a$l_{1}$-norm constraint on the deepest level representations, ignoring the constraints on different level representations. Meanwhile, they fail to discover effectively the essential discrimination information. Therefore, the obtained representations are less discriminative, which degrades model performance. To tackle those issues, we propose an intra- and inter-class induced discriminative deep dictionary learning (DDDL). Specifically, both intra-class compactness and inter-class separability of layer-wise data representations are newly devised as two discriminative constraints on deep dictionary learning. In a hierarchical structure, we obtain a more informative dictionary and the class-specific representations are thus more discriminative at each layer. Due to the$l_{2}$-norm intra- and inter-class constraints of layer-wise data representation, we devise a layer-wise optimization strategy to efficiently learn the closed-form solution of the deepest representation for classification. Comprehensive experiments and analyses on several visual recognition tasks show that our DDDL model surpasses recent shallow and deep representation learning approaches.
Jianping Gou, Xia Yuan, Baosheng Yu, Zhang Yi 0001
IEEE Trans. Multim.2
2022 Fractional Optimization Model for Infrared and Visible Image Fusion
Zhiliang Wu, Xia Yuan, Chunxia Zhao
BMVC4
2022 Deep Dictionary Learning with an Intra-Class Constraint
abstract
In recent years, deep dictionary learning (DDL)has attracted a great amount of attention due to its effectiveness for represen-tation learning and visual recognition. However, most existing methods focus on unsupervised deep dictionary learning, failing to further explore the category information. To make full use of the category information of different samples, we pro-pose a novel deep dictionary learning model with an intra-class constraint (DDLIC) for visual classification. Specif-ically, we design the intra-class compactness constraint on the intermediate representation at different levels to encour-age the intra-class representations to be closer to each other, and eventually the learned representation becomes more dis-criminative. Unlike the traditional DDL methods, during the classification stage, our DDLIC performs a layer-wise greedy optimization in a similar way to the training stage. Experi-mental results on four image datasets show that our method is superior to the state-of-the-art methods.
Xia Yuan, Jianping Gou, Baosheng Yu, Zhang Yi 0001
ICME1
2022 Data Distribution Transfer for Out Of Distribution Generalization
abstract
Modern deep neural networks suffer from performance degradation when evaluated on testing data under different distributions from training data. The goal of out-of-distribution generalization is to solve this problem by learning transferable knowledge from source domains to generalize to invisible target domains. This paper presents a data augmentation method for out-of-distribution generalization. The main assumption is that the main data distribution of an image mostly contains domain-related information, such as color, illumination, texture content, etc, which hurts the domain shifts. To force the model to pay less attention to this part of the information, we propose a new data augmentation method based on the main distribution transition. Extensive experiments on two data set have demonstrated that the proposed method is able to achieve state-of-the-art performance for domain generalization. At the same time, our method can not only combine with other methods to produce a superposition generalization effect but also generate obfuscation data cheaply.
Fawu Wang, Xia Yuan, Chunxia Zhao
MMSP4
2022 Deep Relevant Feature Focusing for Out-of-Distribution Generalization
Fawu Wang, Xia Yuan, Chunxia Zhao
PRCV (1)4
2022 CFNet: Context fusion network for multi-focus images
abstract
Abstract Multi‐focus image fusion aims to generate a clear image by fusing multiple source images. Existing deep learning‐based fusion methods often neglect the context information resulting in the loss of detail information. To address this issue, a context fusion network to merge multi‐focus images, namely CFNet, is proposed. Specifically, a context fusion module is proposed to make full use of low‐level pixels and high‐level semantic features. Particularly, the pyramid fusion mechanism and cross‐scale transfer strategy are adopted to ensure the visual and semantic consistency of the fused image. Meanwhile, to extract salient features more effectively, a spatial attention mechanism is introduced to enhance these features. Further, the pyramid loss is used to progressively refine the fused features at each scale. Experimental results show that the proposed method is superior to some existing methods in both qualitative and quantitative evaluation.
Zhiliang Wu, Xia Yuan, Chunxia Zhao
IET Image Process.3
2022 Competitive binary multi-objective grey wolf optimizer for fast compact antenna topology optimization
abstract
We propose a competitive binary multi-objective grey wolf optimizer (CBMOGWO) to reduce the heavy computational burden of conventional multi-objective antenna topology optimization problems. This method introduces a population competition mechanism to reduce the burden of electromagnetic (EM) simulation and achieve appropriate fitness values. Furthermore, we introduce a function of cosine oscillation to improve the linear convergence factor of the original binary multi-objective grey wolf optimizer (BMOGWO) to achieve a good balance between exploration and exploitation. Then, the optimization performance of CBMOGWO is verified on 12 standard multi-objective test problems (MOTPs) and four multi-objective knapsack problems (MOKPs) by comparison with the original BMOGWO and the traditional binary multi-objective particle swarm optimization (BMOPSO). Finally, the effectiveness of our method in reducing the computational cost is validated by an example of a compact high-isolation dual-band multiple-input multiple-output (MIMO) antenna with high-dimensional mixed design variables and multiple objectives. The experimental results show that CBMOGWO reduces nearly half of the computational cost compared with traditional methods, which indicates that our method is highly efficient for complex antenna topology optimization problems. It provides new ideas for exploring new and unexpected antenna structures based on multi-objective evolutionary algorithms (MOEAs) in a flexible and efficient manner.
Jian Dong 0001, Xia Yuan, Meng Wang 0001
Frontiers Inf. Technol. Electron. Eng.2
2022 Hierarchical Graph Augmented Deep Collaborative Dictionary Learning for Classification
abstract
Recently, deep dictionary learning (DDL) has aroused attention due to its abilities of learning multiple different dictionaries and extracting multi-level abstract feature representations for samples. It has been applied to many intelligent recognition tasks, such as vehicle detection, traffic sign recognition and driver monitoring. Nevertheless, the off-the-shelf DDL-based methods ignore the essential structural information of data in multi-layer dictionary learning. The learned hierarchical data representations are less discriminative. To address this issue, we develop a new DDL framework, called the hierarchical graph augmented deep collaborative dictionary learning (HGDCDL). Firstly, we propose a new deep collaborative dictionary learning (DCDL) that applies collaborative representation to the deepest-level representation learning. Most importantly, equipped with a simple yet effective hierarchal graph construction mechanism, our HGDCDL uses the structure of data to regularize dictionary learning, and generates more informative dictionaries and discriminative representations at different levels. Extensive experiments show that our HGDCDL performs significantly better than the state-of-the-art shallow and deep representation learning methods for classification.
Jianping Gou, Xia Yuan, Lan Du 0002, Shuyin Xia, Zhang Yi 0001
IEEE Trans. Intell. Transp. Syst.2
2021 Salient Object Detection Via Attention-Aware Cascaded Bottom-up Feature Aggregation
abstract
Fully convolutional neural network-based salient object detection has recently achieved great success with its performance benefits from the effective use of multi-layer features. Based on this, most of the existing saliency detectors design complex network structures to fuse the multi-level features of the backbone feature network. However, information in different layer play different roles in saliency object detection, how to integrate them is still an open problem. In this paper, a cascaded bottom-up feature aggregation module is designed to retain and strengthen more spatial details in the low-level features, and embed attention mechanism in the process of feature aggregation to filter more effective features. Extensive experiments show that the proposed networks can consistently improve saliency detection performance. The experimental results on five public datasets prove that this network is competitive in saliency detection.
Fengming Sun, Lufei Huang, Xia Yuan, Chunxia Zhao
ICME3
2021 Binary MOGWO Based On Competition and Teaching for Computationally Complex Engineering Applications
abstract
A new efficient optimization method, called binary multi-objective grey wolf optimizer based on competition and teaching (BMOGWO-CT) mechanisms, is proposed for computationally complex engineering applications. The proposed algorithm first divides the population into four parts belonging to three levels through the competition mechanism, thereby reducing the population number during the following procedure of position updating. Then, the teaching mechanism supervises different parts to update their positions according to their priorities within the whole population therefore further reducing the computational cost for solving the problems. To check the effectiveness of the method, the BMOGWO-CT is tested on ten benchmark test functions and compared to other population-based optimization methods, indicating BMOGWO-CT is more effective and efficient. Furthermore, the novel optimization method is extended to a computationally time-consuming engineering problem- multi-objective optimization of antenna topology. This example verifies the effectiveness of our proposed method in engineering applications.
Xia Yuan, Jian Dong 0001, Meng Wang 0001
IJCNN1
2020 Anisotropic Convolutional Networks for 3D Semantic Scene Completion
abstract
As a voxel-wise labeling task, semantic scene completion (SSC) tries to simultaneously infer the occupancy and semantic labels for a scene from a single depth and/or RGB image. The key challenge for SSC is how to effectively take advantage of the 3D context to model various objects or stuffs with severe variations in shapes, layouts, and visibility. To handle such variations, we propose a novel module called anisotropic convolution, which properties with flexibility and power impossible for the competing methods such as standard 3D convolution and some of its variations. In contrast to the standard 3D convolution that is limited to a fixed 3D receptive field, our module is capable of modeling the dimensional anisotropy voxel-wisely. The basic idea is to enable anisotropic 3D receptive field by decomposing a 3D convolution into three consecutive 1D convolutions, and the kernel size for each such 1D convolution is adaptively determined on the fly. By stacking multiple such anisotropic convolution modules, the voxel-wise modeling capability can be further enhanced while maintaining a controllable amount of model parameters. Extensive experiments on two SSC benchmarks, NYU-Depth-v2 and NYUCAD, show the superior performance of the proposed method.
Jie Li 0040, Kai Han 0001, Peng Wang 0023, Yu Liu 0029, Xia Yuan
CVPR5
2020 Real-time keypoints detection for autonomous recovery of the unmanned ground vehicle
abstract
The combination of a small unmanned ground vehicle (UGV) and a large unmanned carrier vehicle allows more flexibility in real applications such as rescue in dangerous scenarios. The autonomous recovery system, which is used to guide the small UGV back to the carrier vehicle, is an essential component to achieve a seamless combination of the two vehicles. This study proposes a novel autonomous recovery framework with a low‐cost monocular vision system to provide accurate positioning and attitude estimation of the UGV during navigation. First, the authors introduce a light‐weight convolutional neural network called UGV‐KPNet to detect the keypoints of the small UGV form the images captured by a monocular camera. UGV‐KPNet is computationally efficient with a small number of parameters and provides pixel‐level accurate keypoints detection results in real‐time. Then, six degrees of freedom (6‐DoF) pose is estimated using the detected keypoints to obtain positioning and attitude information of the UGV. Besides, they are the first to create a large‐scale real‐world keypoints data set of the UGV. The experimental results demonstrate that the proposed system achieves state‐of‐the‐art performance in terms of both accuracy and speed on UGV keypoint detection, and can further boost the 6‐DoF pose estimation for the UGV.
Jie Li 0040, Kai Han 0001, Xia Yuan, Chunxia Zhao, Yu Liu 0029
IET Image Process.4
2019 RGBD Based Dimensional Decomposition Residual Network for 3D Semantic Scene Completion
abstract
RGB images differentiate from depth as they carry more details about the color and texture information, which can be utilized as a vital complement to depth for boosting the performance of 3D semantic scene completion (SSC). SSC is composed of 3D shape completion (SC) and semantic scene labeling while most of the existing approaches use depth as the sole input which causes the performance bottleneck. Moreover, the state-of-the-art methods employ 3D CNNs which have cumbersome networks and tremendous parameters. We introduce a light-weight Dimensional Decomposition Residual network (DDR) for 3D dense prediction tasks. The novel factorized convolution layer is effective for reducing the network parameters, and the proposed multi-scale fusion mechanism for depth and color image can improve the completion and segmentation accuracy simultaneously. Our method demonstrates excellent performance on two public datasets. Compared with the latest method SSCNet, we achieve 5.9% gains in SC-IoU and 5.7% gains in SSC-IOU, albeit with only 21% network parameters and 16.6% FLOPs employed compared with that of SSCNet.
Jie Li 0040, Yu Liu 0029, Dong Gong, Qinfeng Shi, Xia Yuan, Chunxia Zhao, Ian D. Reid 0001
CVPR5
2019 Histograms of the Normalized Inverse Depth and Line Scanning for Urban Road Detection
abstract
In this paper, we propose to fuse the geometric information of a 3-D LiDAR and a monocular camera to detect the urban road region ahead of an autonomous vehicle. Our method takes advantage of both the high definition of 3-D LiDAR data and the continuity of road in image representation. First, we obtain an efficient representation of LiDAR data and an organized 2-D inverse depth map, by projecting the 3-D LiDAR points onto the camera's image plane. Through the new representation, we can acquire the intermediate representations of road scenes by extracting the vertical and horizontal histograms of the normalized inverse depth. The approximate road regions can be quickly estimated with both histogram-based schemes. To accurately find the road area, we propose a row and column scanning strategy in the approximate road region to refine the detected road area. We have carried out experiments on the public KITTI-Road benchmark, and have achieved one of the best performances among the LiDAR-based road detection methods without learning procedure.
Shuo Gu, Yigong Zhang, Xia Yuan, Jian Yang 0003, Tao Wu 0001, Hui Kong 0001
IEEE Trans. Intell. Transp. Syst.3
2018 Least squares twin bounded support vector machines based on L1-norm distance metric for classification
Qiaolin Ye, Tian'an Zhang, Dongjun Yu, Xia Yuan, Yiqing Xu, Liyong Fu
Pattern Recognit.5
2011 Automatic Segmentation of Head-and-Shoulder Images by Combining Edge Feature and Shape Prior
abstract
Automatic segmentation without any user interaction is very difficult due to potentially high complexity of the scene. No wonder, most existing segmentation algorithms are based on user interactions. However, automatic segmentation in some special situations has great significance. In this paper, we introduce an automatic segmentation algorithm for frontal head-and-shoulder images. Our algorithm combines edge feature and shape prior to extract the foreground silhouette automatically. The novelty of our approach lies in two aspects, namely, the Cost Path Segmentation (CPS) algorithm to extract the initial foreground silhouette, and a general active prior shape model, to extract the final foreground segmentation. We demonstrate the high quality and performance of the proposed approach with a variety of head-and-shoulder images. Compared with previous methods, our approach is much more robust for images with complex color distributions in foreground and background.
Xia Yuan, Fan Zhong 0001, Yijiang Zhang, Qunsheng Peng 0001
CAD/Graphics1
2008 Road-surface abstraction using ladar sensing
abstract
Propose a road-surface abstraction algorithm which suitable for structured and semi-structured road environments. Algorithm uses fuzzy cluster method which based on maximum entropy theory to cluster ladar points that belong to a scan line. After fitting clustered data linearly, one can abstract straight lines that belong to road-surface by their location and slope angle. We can acquire a current referenced horizontal by comparing several continuous ladar scan lines and then the algorithm abstracts obstacles on road-surface area. Experiments show our algorithm works well in spite of the road-boundary's shape is regular or not, and free from the impact of complex texture or irregular illumination of the road.
Xia Yuan, Chunxia Zhao, Yun-fei Cai, Haofeng Zhang 0001, Debao Chen
ICARCV1